Skip to main content
Regression TestingTest Suite OptimizationCI/CD Strategy

How to Build a Regression Suite That Doesn't Slow Down Releases

2 September 2026 · OpenCrevo

Your regression suite started as 40 tests covering the features that mattered most. Three years later it's 800 tests, takes 2 hours to run, and nobody on the team can tell you which of those 800 tests would actually catch a real regression versus which ones are re-testing the same code path a different way. Most existing advice on regression suite optimization comes from academic papers on test-case prioritization algorithms, not from teams who've actually pruned a real suite under release pressure. This article is the practical version: how to classify what you have, cut what's redundant, and structure what's left so it runs fast without losing the coverage that actually matters.

The Real Problem

A regression suite that grows by "add a test for every bug fix and every new feature" without ever removing anything eventually hits a point where its size actively works against its purpose. Two symptoms tell you you're there: releases get delayed waiting for a 2-hour suite to finish, and when it fails, nobody trusts the failure enough to block the release on it, because half the time it's not a real regression, it's redundant coverage that broke for an unrelated reason.

At that point, the suite isn't protecting the release. It's a tax on it.

Why This Happens

Regression suites grow by accretion, not by design. Every bug fix adds a test, which is good practice in isolation, but nobody ever revisits whether the 12 existing tests covering adjacent code paths are still all necessary once the 13th is added. Coverage becomes additive rather than curated. Combined with a common but flawed instinct, that deleting a test feels risky even when it's redundant, because nobody wants to be the person who removed the one test that would have caught the next bug, the suite only ever grows.

The result is a suite where test count and actual regression-catching power diverge over time: more tests, not meaningfully more protection.

How to Diagnose It

Before optimizing, get real data on what you have:

  • Run your suite with a code-coverage tool that reports per-test coverage overlap, not just aggregate suite coverage. Tools like `nyc` (Istanbul) for JavaScript, or `coverage.py` with per-test granularity for Python, can show which tests execute nearly identical code paths.
  • Pull your last 6 to 12 months of production incidents and ask, for each one: did an existing regression test cover this code path, and if so, why didn't it catch the bug? This tells you where your suite has real blind spots versus where it has redundant coverage of already-safe paths.
  • Track flake rate per test over the same period. A test that fails intermittently for reasons unrelated to real regressions is actively eroding trust in the whole suite (see the related article on CI pipelines nobody trusts), independent of whether it's also redundant.

Common Approaches That Fail

  • Prioritization algorithms without a pruning step. Academic test-case-prioritization research (which dominates the existing content on this topic) focuses heavily on ordering tests to fail fast, which helps CI feedback speed but doesn't address a suite that's simply too large for what it protects.
  • "Just parallelize it." Sharding a bloated suite across more CI runners reduces wall-clock time but not cost, flake surface area, or maintenance burden. It treats the symptom, not the redundancy.
  • A one-time cleanup sprint. Pruning a suite once and then returning to the same "add a test per fix, never remove" habit just resets the clock on the same problem.

Practical Solution

Classify every test in the suite into one of three tiers, then structure your CI run and your maintenance policy around that classification, not around the suite as one undifferentiated block:

  • Tier 1, Critical path: covers a code path where a regression would be a significant, customer-visible incident (checkout, auth, core data integrity). Runs on every PR.
  • Tier 2, Important but not critical: covers real functionality but a regression here is annoying, not severe. Runs on merge to main, not on every PR.
  • Tier 3, Redundant or low-value: covers a code path already covered by a Tier 1 or Tier 2 test in a materially similar way, or tests an implementation detail rather than an observable behavior. Candidate for deletion, not for a slower CI tier.

This isn't a one-time exercise. Bake the classification into your definition of done for new tests: a new test added for a bug fix gets tagged at creation time, and any test that hasn't failed independently (caught a real issue on its own, not alongside 5 others) in 12 months gets flagged for review, not automatically deleted, but reviewed.

Implementation

Tagging tests by tier in Playwright, using tags that also drive selective CI execution:

import { test } from "@playwright/test";

test("checkout completes with valid discount code @tier1", async ({ page }) => {
  // critical path, runs on every PR
});

test("account settings page renders saved timezone @tier2", async ({ page }) => {
  // important but not release-blocking on every PR
});

A GitHub Actions workflow that runs only Tier 1 on every PR, and the full suite on merge to main:

name: regression-suite
on:
 pull_request:
 push:
 branches: [main]
jobs:
 tier1-on-pr:
 if: github.event_name == 'pull_request'
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx playwright test --grep @tier1
 full-suite-on-main:
 if: github.event_name == 'push'
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx playwright test

Automation

Once tests are tagged, flake rate and per-test coverage-overlap data (from the diagnosis step) can drive an automated quarterly report: which Tier 3 tests haven't independently caught anything in the last review period, which Tier 1 tests have the highest flake rate and need stabilization first, and which production incidents in the period had no corresponding test at any tier, the real coverage gaps, as opposed to redundant coverage.

AI Considerations

An AI agent is well suited to the mechanical parts of this: scanning a suite for near-duplicate assertions across test files, or drafting the tier classification for a first pass based on which source files a test exercises. It is not well suited to the judgment call of "is this redundancy acceptable given our risk tolerance," which should stay a human decision, informed by the incident-history data above rather than by the AI's guess at business risk.

OpenEvident

For Playwright suites specifically, `vindicate-actions`, published under OpenEvident, provides composite GitHub Actions for running Playwright tests in CI with sharded execution and a merged, native job summary, built for projects already scaffolded with CrevoAI. If you're restructuring CI execution around the tiering approach above, that's a directly relevant building block rather than something to reimplement from scratch. `[VERIFY API BEFORE PUBLICATION]`: the exact action inputs for combining sharding with tag-based test selection (`--grep`) as shown in this article; confirm against the current `vindicate-actions` README before using in production.

OpenCrevo Implementation

Restructuring a suite that's grown for years without a tiering discipline is real, ongoing work, not a weekend task, especially alongside a normal release schedule. OpenCrevo's Test Automation service builds repeatable, CI-integrated test suites designed with this kind of tiering from the start; if your team needs this done alongside continued feature delivery rather than during a dedicated freeze, OpenCrevo can help design and implement the migration. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Classify every existing test into Tier 1, Tier 2, or Tier 3 before changing your CI configuration.
  • Run Tier 1 only on every PR; move Tier 2 and Tier 3 to a less frequent schedule.
  • Cross-reference the last 6 to 12 months of production incidents against your suite to find real coverage gaps, not just redundancy.
  • Track per-test flake rate separately from redundancy; a flaky Tier 1 test needs stabilization, not deletion.
  • Bake tier classification into the definition of done for every new test going forward, not just as a one-time audit.

FAQ

  • Won't moving tests out of the per-PR run reduce coverage on every change? Not if Tier 1 genuinely covers the critical paths; Tier 2 and Tier 3 still run on merge to main, so regressions are caught before release, just not blocking every individual PR.
  • How do we decide what counts as "critical path" for Tier 1? Start from the incident-history review: any code path that has caused a significant production incident in the last 12 to 24 months is Tier 1 by default; expand from there based on business risk, not test-writer preference.
  • Is this the same as test-case prioritization from academic research? Related but not the same. Prioritization orders tests to fail fast within a run; this framework decides which tests run at all in a given CI trigger, and which get pruned entirely.
  • How often should the tiering be reviewed? Quarterly is a reasonable cadence for most teams, `(Opinion)`, more frequently for suites growing quickly or shipping at high velocity.

Conclusion

A regression suite that's too slow to trust isn't a sign you need more automation, it's a sign the suite has grown without a maintenance discipline to match its growth. Classifying tests by real risk, pruning genuine redundancy, and structuring CI execution around that classification gets you a suite that's both faster and more trustworthy, not a trade-off between the two.

Sources

Playwright documentation on test tagging and grep-based selection, GitHub Actions composite actions documentation, OpenEvident's `vindicate-actions` repository, OpenCrevo services documentation (`src/data/services.ts` in this repository). Note per `content-gap-analysis.md`: most existing published research on this exact topic is academic (IEEE, PLOS) rather than practitioner-authored; this article intentionally departs from that format.

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.