Skip to main content
PR Test SelectionSmart Test SelectionCI Test Strategy

How to Decide Which Tests Should Run on Every Pull Request

2 September 2026 · OpenCrevo

Every team with more than a handful of tests eventually hits the same wall: the full suite takes too long to run on every pull request, but nobody wants to be the one who decides which tests get skipped. Most existing content on this problem comes from a single vendor, Codecov, framing it as a coverage-diff problem their product solves. That framing is incomplete: test selection is a risk-allocation decision, not a coverage-math problem, and it can be implemented with plain CI configuration you already have access to. This article gives you a vendor-neutral decision framework, plus a real GitHub Actions path-filter example and a tag-based selection example, so you can implement PR-level test selection without adopting a new product to do it.

The Real Problem

A pull request triggers CI. CI runs the full test suite. If the suite takes 5 minutes, nobody minds. If it takes 45 minutes or 2 hours, one of two things starts happening: engineers stop waiting for it and merge anyway once a manual smoke check looks fine, or the team quietly starts skipping CI on "small" changes, which is exactly the kind of change most likely to have an unreviewed side effect.

Neither outcome is a testing failure exactly. It's a scheduling failure: every test is being asked to prove the same thing on every change, regardless of whether that change could plausibly affect what the test verifies. A one-line CSS tweak in a marketing footer component doesn't need the payment-processing integration suite to run before merge. A change to the checkout service does.

The question this article answers is not "how do we make tests faster" (that's a separate, valid problem). It's: given a specific diff, which subset of your existing tests should gate that specific PR, and which can safely run later.

Why This Happens

Most teams don't design test-selection policy up front, because early on, the full suite is fast enough that there's no policy to design. The suite runs everything, every time, and that's fine until it isn't. By the time it becomes a real problem, the team has hundreds of tests with no metadata attached to them describing what code paths they actually exercise or how critical those paths are `(Industry consensus)`.

At that point, two forces push toward inaction. First, nobody wants to be the person who excludes a test from the PR gate and later gets blamed when a bug that test would have caught reaches production, even though that same test running 45 minutes late on every single PR is also a real cost, just a diffuse and less visible one. Second, path-based or tag-based selection requires actually mapping tests to the code they cover, which is unglamorous, one-time setup work that competes for time against feature delivery.

The result is a suite that either runs everything on every PR (slow, but "safe") or gets arbitrarily thinned by whoever is annoyed enough that week (fast, but unprincipled). Neither is a policy. This is also why the space is dominated by coverage-diff products: they sell a way to skip the mapping work by inferring it from code coverage data automatically, which is a reasonable product to buy, but it isn't the only way to solve the underlying problem, and understanding the underlying decision logic matters even if you do eventually buy a tool for it.

Common Approaches That Fail

  • Run everything, every time. Simple to reason about, but doesn't scale past a suite that fits comfortably inside a coffee break. Teams tolerate this far longer than they should because nobody owns the decision to change it.
  • Coverage-diff tooling as the only strategy. Products like Codecov's test-selection features infer which tests to run from code coverage overlap with the diff. This works, but it's a black box from the team's perspective: when the tool skips a test incorrectly, engineers can't reason about why without opening the vendor dashboard, and it's a paid dependency for a decision your own CI config can express directly for a large share of cases `(Opinion)`.
  • "Just run unit tests on PR, everything else on merge." A test-type split (unit versus integration versus E2E) is not the same axis as risk. A unit test for a critical billing calculation is more important to gate a PR than an E2E test for a rarely used settings page, but a pure test-type split treats all unit tests as equally PR-worthy and all E2E tests as equally deferrable.
  • Manual judgment call per PR. Asking the PR author or reviewer to decide "does this need the full suite" reintroduces exactly the inconsistency and social pressure that caused the problem: nobody wants to be the one who says "skip it" and be wrong.

Practical Solution

Test selection on a PR is really two independent decisions layered together, and treating them as one is where most ad hoc approaches go wrong:

  • What changed? (path-based selection: which parts of the codebase does this diff actually touch)
  • How important is what changed? (risk-tier selection: of the tests that cover the touched area, which are critical enough to block merge versus safe to defer)

This builds directly on the Tier 1/Tier 2/Tier 3 risk classification from regression suite pruning: that article covers how to classify tests by risk once you already know which ones are candidates to run. This article covers the layer above it: how to decide which tests are even candidates for a given PR, based on what changed, before tier logic ever applies.

The combined framework:

  • Step 1, path filter. Use your CI's path-filtering capability to determine which service, package, or directory the PR touches. A PR that only touches `docs/` or `marketing-site/` shouldn't trigger the backend test suite at all.
  • Step 2, tag filter within the touched area. Within whatever test suite corresponds to the touched path, run only the tests tagged Tier 1 (critical path, per the regression-suite tiering) on the PR itself. Defer Tier 2 and Tier 3 to the merge-to-main run.
  • Step 3, an explicit escape hatch. Any PR can be labeled to force the full suite (a `run-full-suite` label, or similar), for the cases where path/tag inference genuinely can't judge risk, a shared utility function touched by everything, a dependency bump, a config change. This keeps the system from becoming a black box that engineers fight against; it stays inspectable and overridable.

This is the concrete "vendor-neutral decision framework with real code examples" this space is missing: a coverage-diff product infers step 1 and 2 from historical coverage data automatically, which is a legitimate shortcut, but the underlying logic, path relevance plus risk tier, is the same decision either way, and you can implement it in plain CI configuration first and evaluate whether you need a paid layer on top later.

Implementation

Path-filter selection (GitHub Actions, using `dorny/paths-filter`)

This determines which service's test suite should even run, based on what the PR actually touches:

name: pr-test-selection
on:
 pull_request:

jobs:
 detect-changes:
 runs-on: ubuntu-latest
 outputs:
 backend: ${{ steps.filter.outputs.backend }}
 frontend: ${{ steps.filter.outputs.frontend }}
 steps:
      - uses: actions/checkout@v4
      - uses: dorny/paths-filter@v3
 id: filter
 with:
 filters: |
 backend:
              - 'services/payments/**'
              - 'services/checkout/**'
              - 'shared/lib/**'
 frontend:
              - 'apps/web/**'
              - 'packages/ui/**'

 backend-tier1:
 needs: detect-changes
 if: needs.detect-changes.outputs.backend == 'true'
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx playwright test --grep @tier1 --project=backend

 frontend-tier1:
 needs: detect-changes
 if: needs.detect-changes.outputs.frontend == 'true'
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx playwright test --grep @tier1 --project=frontend

Note `shared/lib/**` is deliberately mapped to the backend filter here, not excluded: a change to shared code is exactly the case where path-based inference alone under-selects, which is why the escape hatch in the next example matters.

Tag-based selection plus the escape-hatch label

Building on the `@tier1` / `@tier2` tagging convention from the regression-suite article, add an explicit override for PRs where the author knows path/tag inference won't judge risk correctly:

 full-suite-override:
 if: contains(github.event.pull_request.labels.*.name, 'run-full-suite')
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx playwright test
import { test } from "@playwright/test";

test("checkout completes with valid discount code @tier1", async ({ page }) => {
  // path: services/checkout, risk: critical, runs on every PR touching checkout
});

test("shared currency formatter handles negative amounts @tier1", async ({ page }) => {
  // path: shared/lib, risk: critical because it's used everywhere,
  // this is why shared/lib maps to the backend PR filter above, not a separate low-priority lane
});

The combination of the two steps above is deliberately simple to read: a PR touching only `apps/web/` never triggers backend tests, a PR touching `shared/lib/` triggers the full backend Tier 1 set because shared code is treated as high-risk by default, and any PR can force the complete suite with one label when the author has reason to believe the automatic selection isn't enough. That inspectability, an engineer can read the YAML and know exactly why a given test did or didn't run, is the actual gap in a purely coverage-diff-driven approach: those tools work, but the selection logic lives inside a third-party service rather than in a file the team owns and can read.

AI Considerations

An AI coding agent is genuinely useful for the setup cost of this framework: mapping existing tests to the paths they cover so path filters and tags can be added retroactively, since that mapping is exactly the kind of large, mechanical, first-pass classification work an agent can do faster than a human reading through hundreds of test files `(Opinion)`. It's a poor fit for deciding risk tier on its own: whether a code path is "critical" is a business-risk judgment tied to what an incident there would actually cost the company, not something inferable purely from code structure or historical coverage. Treat AI output here as a draft classification for a human to confirm against real incident history, the same caution that applies to AI-generated tests generally, not as the final policy.

OpenEvident

If you're implementing tag-based selection on a Playwright suite, `vindicate-actions`, published under OpenEvident, provides composite GitHub Actions for running CrevoAI-scaffolded Playwright tests in CI with sharded execution and a native job summary, which pairs directly with the `--grep @tier1` pattern shown above once a suite is already scaffolded with CrevoAI. `[VERIFY API BEFORE PUBLICATION]`: whether `vindicate-actions`' `run-shard` action accepts a `--grep` pattern or an equivalent tag filter as an input alongside its sharding logic; confirm against the current README before using this combination in production.

OpenCrevo Implementation

Retrofitting path-and-tag-based selection onto a suite that's been running everything on every PR for years is real implementation work: mapping existing tests to code paths, agreeing on risk tiers across teams that may not currently agree on what's "critical," and rewriting CI configuration without breaking the release gate mid-migration. OpenCrevo's Test Automation service builds repeatable, CI-integrated test suites designed with this kind of selective execution from the start; if your team needs the retrofit done alongside continued feature delivery, OpenCrevo can help design and implement it. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Map existing tests to the code paths they cover before writing any CI selection logic; you cannot filter by path without this mapping.
  • Layer path filtering (what changed) and risk-tier filtering (how important is it) as two separate, composable steps, not one combined rule.
  • Add an explicit, labeled escape hatch to force the full suite for PRs where automatic selection can't be trusted, shared utilities, dependency bumps, config changes.
  • Keep the selection logic in version-controlled CI configuration your team can read, rather than depending solely on an opaque third-party inference engine.
  • Revisit the path-to-test mapping whenever the codebase is restructured; a stale mapping silently under-selects tests for moved code.

FAQ

  • Is this worth it for a small team with a fast suite? Not yet. If your full suite runs in a few minutes, the setup cost of path and tag mapping isn't worth it `(Opinion)`. Revisit once suite time starts affecting how often people wait for CI before merging.
  • How long does this take to implement? For a suite that's already tagged by tier (per the regression-suite classification), adding path filters is a config change measured in hours. Building the tier tagging from scratch first, if it doesn't exist yet, is the larger, ongoing effort.
  • Does this replace coverage-diff tools like Codecov entirely? Not necessarily. This framework covers the majority of cases with plain CI config; a coverage-diff product can still add value for the harder cases, code with unclear or undocumented path-to-test mapping, at the cost of an external dependency and less inspectable selection logic.
  • What's the difference between this and the regression-suite tiering article? That article defines the tiers (Tier 1/2/3) and decides overall CI cadence per tier. This article decides, for a specific PR's diff, which tests are even candidates before tier logic applies.
  • What breaks first when a team adopts this at scale? Usually the path-to-test mapping goes stale after a codebase restructuring, silently under-selecting tests for moved or renamed code, `(Industry consensus)`, which is why the mapping needs an owner, not just a one-time setup task.

Conclusion

Deciding which tests run on every PR isn't a coverage-math problem you need a vendor dashboard to solve; it's a two-layer decision, what changed and how important is it, that plain path filtering and tag-based selection in your existing CI config can express directly and transparently. Start with the mapping work, layer risk tiers on top, keep an explicit override for the cases inference can't judge, and you get a PR gate that's both fast and inspectable, without handing the decision logic to a third party.

Sources

GitHub Actions documentation on pull_request triggers and conditional jobs, dorny/paths-filter GitHub Action documentation, Playwright documentation on test tagging and grep-based selection, OpenEvident's vindicate-actions repository, OpenCrevo services documentation (src/data/services.ts in this repository). Note per content-gap-analysis.md: existing top-ranking content in this specific niche is largely product-led (Codecov); this article intentionally supplies the underlying vendor-neutral logic instead.

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.