Skip to main content
Test TriageFlaky TestsCI Automation

Test Failure or Product Bug? A Practical Failure-Triage Workflow

2 September 2026 · OpenCrevo

A test fails in CI. Is it a real product bug, a flaky test, a stale test asserting outdated behavior, or an environment problem? Most teams answer this question by gut feel, and the answer usually depends on who's looking at it and how much time they have, which means the same failure gets triaged differently depending on the day. Generic bug-triage content doesn't help here, because it's written for bugs reported by users, not for automated test failures, which need a different first question: is this even a real defect at all? This article gives a concrete, repeatable triage flowchart built specifically for that first question.

The Real Problem

Your CI pipeline has 200 tests and 3 failed this run. A developer looks at the failures, sees a stack trace that doesn't immediately make sense, shrugs, and re-runs the pipeline. It passes the second time. Nobody investigates further. Three of those "flaky, ignore it" dismissals over the following month turn out to have been the same real, intermittent race condition in the checkout service, one that eventually caused a customer-facing incident. The team didn't have a bad test suite; they had no consistent process for deciding when a red check deserved investigation versus a re-run.

Why This Happens

Automated test failures collapse four genuinely different situations into one signal: a red check. Those four situations need different responses, but nothing in a typical CI dashboard distinguishes them:

  • A real product regression. The code changed, the behavior is now actually wrong.
  • A flaky test. The test itself has a timing, ordering, or environment dependency unrelated to the actual feature it's testing.
  • A stale test. The product's intended behavior changed (a legitimate requirement update), but the test wasn't updated to match, so it's failing against a now-outdated expectation.
  • An environment or infrastructure problem. The test and the product are both fine; a dependency, a test database, or a CI runner issue caused the failure.

Without a structured way to tell these apart quickly, teams default to the path of least resistance: re-run and move on. That's a reasonable response to genuine flakiness, and a dangerous one for a real regression.

How to Diagnose It

Before building the triage workflow, get a baseline: pull the last month of CI failures and manually classify each one into the four categories above. This tells you two things: your team's actual flake-versus-real-bug ratio (useful context for how much triage rigor is warranted), and which specific tests or code areas produce the most ambiguous failures, a good starting point for the workflow's first decision point.

Common Approaches That Fail

  • "Just re-run it." Correct response to genuine flakiness, dangerous default response to everything, because it never distinguishes flakiness from a real regression before deciding to ignore the signal.
  • Generic bug-triage severity matrices. Standard bug-triage frameworks (severity, priority, assignee) assume you already know it's a real bug. They don't help with the automation-specific first question of whether it's a real bug at all.
  • Requiring a human to investigate every single failure in depth. Doesn't scale past a small suite, and burns out whoever's on triage duty, which leads right back to the "just re-run it" default.

Practical Solution

A four-question triage flowchart, applied to every CI failure before any deeper investigation:

  • Question 1: Did this test pass on a re-run with no code changes? Yes, it passed on re-run: proceed to Question 2 (flakiness suspected, but confirm before dismissing). No, it failed consistently: proceed to Question 3 (likely real, skip to deeper investigation).
  • Question 2 (only reached if re-run passed): Has this specific test failed intermittently before? Yes, check your flake-tracking log (see Implementation below): confirm it's a known flaky test, tag the failure as flaky, and separately track it for stabilization work, don't just ignore it. No, this is the first time: treat it as a genuine unknown, not confirmed flakiness. Investigate once before classifying it as flaky, since a first-time intermittent failure is exactly how the checkout race condition in the scenario above got dismissed.
  • Question 3 (reached if it failed consistently): Does the test's assertion match the current, intended product behavior? Yes, the assertion is correct and the product now violates it: this is a real regression. Escalate immediately, block the release if it's Tier 1 (see the related regression-suite article for tiering). No, the product behavior legitimately changed and the test wasn't updated: this is a stale test. Update the test's assertion to match the new intended behavior, and note in the PR why.
  • Question 4 (a branch reachable from either path): Is the failure isolated to one environment or CI runner? Yes: investigate infrastructure (dependency version drift, test database state, runner resource limits) before touching the test or the product code at all.

Implementation

A minimal flake-tracking log, as a simple structured record rather than tribal knowledge, keeps Question 2 answerable in seconds instead of requiring someone to remember:

{
  "test_name": "checkout completes with valid discount code",
  "first_seen_flaky": "2026-06-14",
  "occurrences": 4,
  "last_occurrence": "2026-08-20",
  "status": "known_flaky_investigating",
  "linked_issue": "JIRA-4821"
}

This can live as a simple JSON or database table updated by a CI post-failure hook, or as a dedicated field in whatever issue tracker the team already uses; the mechanism matters less than having it be queryable in seconds during triage, not something someone has to remember.

A lightweight GitHub Actions step that auto-reruns a failed job once and flags the result for the triage log, rather than silently retrying and hiding the signal:

- name: Run tests
 id: first-run
 run: npx playwright test
 continue-on-error: true
- name: Retry on failure
 if: steps.first-run.outcome == 'failure'
 run: npx playwright test
- name: Flag for flake log if retry passed
 if: steps.first-run.outcome == 'failure'
 run: echo "Failed then passed on retry, log for flake tracking" >> $GITHUB_STEP_SUMMARY

Automation

Once the flake-tracking log exists, the "did this fail before" check in Question 2 can be automated entirely: a bot comment on the PR or CI summary that queries the log and states whether this specific test has a known-flaky history, removing the manual lookup step from triage and making the workflow fast enough to apply consistently rather than only when someone has time.

AI Considerations

An AI agent can reasonably handle the mechanical parts of this flowchart: checking the flake log, re-running a failed job, and drafting a stale-test assertion update once a human confirms the product behavior genuinely changed. The judgment call in Question 3, whether a failing assertion reflects a real regression or a legitimate, intended behavior change, requires understanding the product requirement, not just the code, and should stay a human decision informed by the actual requirement or ticket, not inferred by the AI from the diff alone.

OpenCrevo Implementation

Building this kind of triage discipline into a team's actual CI workflow, including the flake-tracking log and the automated retry-and-flag step, is exactly the kind of process design OpenCrevo's Quality Engineering service delivers: production-grade quality frameworks and test automation built to catch regressions before they reach production, not just individually written tests. If your team has the CI pipeline but not the triage process around it, OpenCrevo can help design and implement it. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Never classify a failure as flaky based on a single re-run alone; check the flake-tracking log for prior occurrences first.
  • Treat a first-time intermittent failure as an unknown requiring one real investigation, not an automatic dismissal.
  • For consistent failures, always ask whether the assertion still matches intended product behavior before assuming it's a real bug.
  • Log every confirmed-flaky test with an occurrence count and a linked stabilization ticket, don't just silently re-run and move on.
  • Separate "infrastructure caused this" from both flakiness and real bugs; it needs a different owner (platform/DevOps) than either.

FAQ

  • Does this replace a regular bug-triage process? No, it's a pre-filter specifically for automated test failures, to decide whether a standard bug-triage process (severity, priority, assignee) should even be invoked for this failure.
  • How much does this slow down responding to a CI failure? The flowchart itself takes seconds once the flake log exists; the time cost is in the one-time setup of the log and the retry automation, not in ongoing use.
  • What if we don't have a flake-tracking log yet? Start one today, even a simple spreadsheet, with just test name, date, and occurrence count. It gets more useful every week; the value isn't in having a complete history on day one.
  • Should every consistently-failing test block a release? Only Tier 1 (critical-path) tests, per the tiering approach in the related regression-suite article; a consistently-failing Tier 3 test still needs investigation, but shouldn't automatically block shipping.

Conclusion

A red CI check is a signal, not a verdict. Treating every failure the same, whether by always re-running or always escalating, either hides real regressions or burns out whoever's on triage duty. A short, consistent flowchart that asks the right four questions in order turns "is this real?" from a gut-feel guess into a repeatable process, which is what actually makes a regression suite worth trusting.

Sources

Playwright documentation on retries and test annotations, GitHub Actions continue-on-error and step summary documentation, this program's own content-gap-analysis.md (identifying this as a genuinely underserved topic relative to generic bug-triage content), OpenCrevo services documentation (src/data/services.ts in this repository).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.