Skip to main content
Test CoverageMutation TestingSoftware Quality

Why Increasing Test Coverage Doesn't Always Improve Software Quality

2 September 2026 · OpenCrevo

A team spends a quarter pushing test coverage from 68 percent to 91 percent. Leadership is pleased. The next release still ships two production incidents that a customer finds before QA does. Nobody on the team is lying about the coverage number, and nobody skipped the work. The number is real. It's also not measuring the thing everyone assumed it was measuring. Coverage tells you which lines of code executed during a test run. It does not tell you whether anything meaningful was verified when they did. This article covers why that gap exists, why it shows up even in teams that aren't cutting corners, and what an actual dashboard redesign looks like when a team stops reporting raw coverage percentage to leadership and starts reporting something leadership can act on.

The Real Problem

Take a payment-retry service: when a card charge fails, the service decides whether to retry immediately, retry with backoff, or give up and notify the customer. A team inherits this service with 40 percent test coverage and is told to get it above 85 percent before the next audit. They add tests. Six weeks later coverage sits at 89 percent. Every new test calls the retry function, passes in a failed-charge object, and asserts the function returned without throwing.

The coverage report is accurate: those lines do execute. What the report can't show is that not one of the new tests checks which retry path was chosen, how many attempts happened, or whether the customer got notified on the give-up path. A defect where "give up" silently retries forever, burning through a customer's card multiple times, would execute every one of those covered lines and still pass every one of those tests. The bug ships. The coverage number goes up regardless.

This is not a story about a careless team. It's what happens by default when coverage percentage becomes the target instead of a byproduct: `(Industry consensus)` a metric that's used as a target tends to get optimized directly, and line/branch coverage is trivially satisfiable by executing code without asserting anything specific about its behavior.

Why This Happens

Coverage percentage measures execution, not verification, for a structural reason: a coverage tool instruments the code and records which lines ran during the test suite. It has no way to inspect whether the assertions in that test actually constrain the code's behavior. A test with `expect(result).not.toBeNull()` and a test with `expect(result.retryCount).toBe(2)` count identically toward branch coverage, even though only the second one would catch a bug in the retry-count logic.

Three specific dynamics push coverage upward without pushing defect-catching power upward alongside it:

  • Coverage is easiest to raise on the paths that are already correct. Adding a test for the straightforward success path is fast and safely passes; adding a test for the interaction between two edge cases (two failed charges in the same billing cycle, a retry that itself times out) takes longer to design and is more likely to reveal an actual bug, which creates a subtle incentive to write the easy test first, and often only the easy test, when the goal is a coverage number by a deadline.
  • A percentage target rewards volume over precision. If the ask is "get to 85 percent," the fastest route is however many shallow tests it takes to touch the remaining lines, not fewer, stronger tests that exercise the business logic those lines represent.
  • Coverage tools can't distinguish a test that exercises logic from a test that merely calls it. This is the same gap mutation testing exists to close, and it's genuinely under-visible until a team looks for it directly, because a coverage report and a green CI pipeline both look identical whether the underlying tests are strong or hollow.

Common Approaches That Fail

  • Setting a hard coverage percentage as a release gate. This produces exactly the incentive described above: the fastest way to clear the gate is the shallowest test that touches the required lines, not the strongest one.
  • Auditing test quality by re-reading every test manually before each release. This doesn't scale past a small suite, and it puts the burden on human attention at exactly the moment (pre-release crunch) when attention is scarcest.
  • Replacing the coverage percentage with no metric at all. Some teams, on realizing coverage is misleading, stop measuring anything structured and fall back to "the team feels good about this release." That's a worse decision input than a flawed metric, not a better one; the fix is a better metric, not the absence of one.
  • Reporting coverage percentage to leadership as the headline quality indicator. Leadership isn't wrong to want a number. The problem is that coverage percentage answers a question ("did we run this code?") that isn't the question leadership actually cares about ("will this release cause an incident?"). Continuing to report the wrong answer to the right question doesn't get fixed by explaining the metric better; it gets fixed by reporting a different metric (see the related article on quality intelligence versus raw test-case count for the same argument applied one level up, to test suites in general rather than coverage specifically).

Practical Solution

The fix isn't abandoning coverage measurement. It's demoting coverage percentage from headline metric to one input among several, and pairing it with metrics that actually correlate with defect risk. Three additions do most of the work:

  • Mutation score, on the critical-path code specifically, not the whole codebase. Mutation testing tools (Stryker, PIT, mutmut) introduce small deliberate bugs into the source and check whether any test catches them. A payment-retry function with 89 percent line coverage and a 35 percent mutation score is telling you directly: most of the tests execute this code without verifying its behavior, which is the exact failure pattern from the retry-count example above.
  • Defect-escape rate, tracked per release: how many bugs were found in production versus caught before release, broken down by the area of the codebase they came from. This is the actual outcome metric coverage percentage is a weak proxy for. If defect-escape rate is flat or rising while coverage percentage rises, that's the clearest possible signal that the coverage increase isn't buying real protection.
  • Critical-path coverage, reported separately from whole-codebase coverage. A settings page and a payment-retry function are not equally risky to leave under-tested, but a single blended coverage percentage treats them as interchangeable. Naming the 10 to 20 percent of the codebase that would cause a real incident if broken, and tracking coverage and mutation score against just that set, gives leadership a number that maps to actual business risk instead of an average across code that mostly doesn't matter if it's wrong.

Implementation

Here is what an actual dashboard redesign looks like in practice: the same engineering leadership review that used to get one number now gets a small table.

Before, the leadership-facing quality slide:

  • Test coverage: 91%

After, the same review, replaced with a table that separates "did we run the code" from "would we catch a real bug" from "did anything actually reach production":

  • Critical-path coverage (payment, auth, data-write paths): This release 87%, last release 79%, target 90%
  • Mutation score, critical path: This release 61%, last release 58%, target 75%
  • Whole-codebase line coverage: This release 91%, last release 84%, target informational only, not gated
  • Defect-escape rate (bugs found in production / total bugs found): This release 18%, last release 24%, target trending down
  • Flaky test rate (tests that failed and passed on rerun with no code change): This release 6%, last release 9%, target under 5%

The whole-codebase line coverage row is kept, deliberately, but relabeled "informational only." This matters for a specific reason: a sudden drop in headline coverage percentage, with no explanation, tends to trigger an unproductive leadership question in the next meeting. Keeping the number visible but explicitly demoted avoids that reflex while making clear it isn't the number the team is managed against. `(Opinion)`: the specific thresholds above (75 percent mutation score, 90 percent critical-path coverage, under 5 percent flaky rate) are reasonable starting points, not universal targets; the right numbers depend on a team's own defect-escape history and risk tolerance.

Producing the mutation-score row without slowing CI to a crawl means scoping it, not running it against the full suite on every commit:

{
  "mutate": ["src/payments/retry-handler.ts", "src/auth/**/*.ts"],
  "testRunner": "command",
  "commandRunner": { "command": "npx jest --testPathPattern=payments|auth" },
  "thresholds": { "high": 80, "low": 60, "break": 50 }
}

Run this scoped configuration nightly or on merges to main, not on every PR, and feed the resulting mutation score into the dashboard table above as a rolling weekly number rather than a per-commit one. That keeps the signal current without making mutation testing a bottleneck in the PR loop.

AI Considerations

AI coding assistants make the underlying problem worse in one specific, predictable way, and better in another. Worse: an AI agent asked to "increase test coverage" will reliably do exactly that, generating tests shaped to touch uncovered lines rather than to verify business logic, because "raise the coverage number" is a well-specified, easy target for a model to optimize toward, and "verify the retry logic handles two simultaneous failures correctly" requires the kind of domain reasoning a coverage-chasing prompt doesn't ask for. `(Industry consensus)` this pattern, generated tests that satisfy a coverage target without strengthening actual verification, is a widely observed failure mode wherever AI is used to close a coverage gap on a deadline.

Better: the same mutation-testing feedback loop that catches this in human-written tests works on AI-generated ones, and can be handed back to the AI agent directly. A prompt of "here is a mutation that survived: the retry limit check was changed from `<=` to `<` and no test failed, write a test that would catch this specific mutation" produces a materially stronger test than "write more tests for the retry handler," because it gives the model a concrete, falsifiable target instead of an open-ended one. Treating mutation score, not coverage percentage, as the metric an AI agent is asked to improve closes most of the gap described above without adding a new manual review step.

OpenCrevo Implementation

Redesigning what a leadership dashboard reports, and building the underlying mutation-testing and critical-path-tagging pipeline that makes those numbers real rather than aspirational, is the kind of work OpenCrevo's Quality Engineering service is built around: production-grade quality frameworks, test automation, and evaluation pipelines built to be trusted, not just to look green in a status report. If your team already has a coverage number everyone privately distrusts and needs the reporting and the underlying pipeline rebuilt around metrics that actually track defect risk, OpenCrevo can help design and implement that transition without requiring a rewrite of the existing test suite. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Stop using a single whole-codebase coverage percentage as a release gate; keep it visible but explicitly informational.
  • Identify the 10 to 20 percent of the codebase that would cause a real incident if broken, and track coverage and mutation score against that set specifically.
  • Run a scoped mutation-testing pass against critical-path code on a nightly or merge-to-main cadence, not necessarily on every PR.
  • Add defect-escape rate and flaky test rate to the same reporting table, since both correlate with real release risk in a way raw coverage doesn't (see the related article on measuring test automation ROI without fake metrics for how the same reporting discipline applies to ROI claims, not just coverage).
  • When asking an AI agent (or a human) to close a coverage gap, ask for a target mutation score or a specific surviving mutation to fix, not a target coverage percentage.

FAQ

  • Does this mean we should stop tracking coverage percentage entirely? No. Keep tracking it, but stop treating it as the headline quality metric or a hard release gate. It's still a useful, cheap signal for "did we forget to test this file at all," just not for "is this code actually verified."
  • How long does it take to set up mutation testing on an existing codebase? Scoped to a specific critical-path directory as shown above, a working configuration can be running within a day for a team already using a standard test runner. Running it against an entire legacy codebase is a much larger, separate project and usually isn't worth doing all at once.
  • What's the difference between coverage and mutation score? Coverage measures whether a line of code executed during testing. Mutation score measures whether the tests would actually fail if that line's logic were subtly wrong. A codebase can have high coverage and a low mutation score at the same time; that combination is the exact signal this article is about.
  • How do you measure success after making this change? Not by watching the coverage number, since it's explicitly de-emphasized. Track defect-escape rate over 2 to 3 release cycles: if mutation score on critical-path code rises and defect-escape rate falls in the same window, the new reporting is doing its job.
  • What breaks first when a team tries this at scale? Usually CI runtime, if mutation testing is applied unscoped to the whole suite instead of the critical-path subset. Scoping it to a defined critical-path directory list, as shown in the Implementation section, is what keeps this practical past a small pilot.

Conclusion

A coverage percentage going up is not the same claim as quality going up, and treating them as interchangeable is what lets a real production incident ship inside a suite that looks, on paper, better than ever. The fix isn't a more disciplined reading of the same number. It's reporting a different, small set of numbers, mutation score on critical-path code, defect-escape rate, flaky test rate, alongside coverage rather than instead of it, so that the dashboard leadership sees actually tracks the risk they're trying to manage.

Sources

Stryker Mutator documentation, PIT mutation testing documentation, industry discussion of coverage-percentage-versus-mutation-score divergence and coverage-as-a-target incentive effects (`Industry consensus`, synthesized from multiple current QA-engineering publications, not a single primary source), OpenCrevo services documentation (src/data/services.ts in this repository).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.