The Real Problem
Your CI pipeline has 200 tests and 3 failed this run. A developer looks at the failures, sees a stack trace that doesn't immediately make sense, shrugs, and re-runs the pipeline. It passes the second time. Nobody investigates further. Three of those "flaky, ignore it" dismissals over the following month turn out to have been the same real, intermittent race condition in the checkout service, one that eventually caused a customer-facing incident. The team didn't have a bad test suite; they had no consistent process for deciding when a red check deserved investigation versus a re-run.