The Real Problem
A pipeline with a 40% chance of a false-red result on any given run isn't a QA inconvenience, it's a broken instrument. If a smoke alarm cries wolf four times out of ten, people stop evacuating when it goes off, and eventually it goes off during a real fire and nobody moves. CI pipelines fail the same way, just slower and quieter.
Here's the pattern almost every team recognizes once it's named: a merge queue backs up because three unrelated PRs all show a red check on the same intermittent test. Someone reruns the job. It goes green. Nobody files anything, because filing something takes longer than clicking rerun, and rerun usually works. Multiply that by every PR, every week, for a year, and you get a team that has fully stopped reading failure output before clicking rerun. The signal is still being produced. Nobody is receiving it anymore.
The cost isn't abstract. A real regression eventually rides through on the back of "it's probably flaky," because the person looking at the failure has no fast way to tell a genuine flake from a genuine bug (Industry consensus: this exact ambiguity is why teams need a structured triage step rather than gut feel, covered in depth in "Test Failure or Product Bug? A Practical Failure-Triage Workflow). By the time someone notices in production, the CI run that should have caught it is three "just rerun it" clicks in the past.