The Real Problem
Take a payment-retry service: when a card charge fails, the service decides whether to retry immediately, retry with backoff, or give up and notify the customer. A team inherits this service with 40 percent test coverage and is told to get it above 85 percent before the next audit. They add tests. Six weeks later coverage sits at 89 percent. Every new test calls the retry function, passes in a failed-charge object, and asserts the function returned without throwing.
The coverage report is accurate: those lines do execute. What the report can't show is that not one of the new tests checks which retry path was chosen, how many attempts happened, or whether the customer got notified on the give-up path. A defect where "give up" silently retries forever, burning through a customer's card multiple times, would execute every one of those covered lines and still pass every one of those tests. The bug ships. The coverage number goes up regardless.
This is not a story about a careless team. It's what happens by default when coverage percentage becomes the target instead of a byproduct: `(Industry consensus)` a metric that's used as a target tends to get optimized directly, and line/branch coverage is trivially satisfiable by executing code without asserting anything specific about its behavior.