Use a formula with four terms instead of one: execution savings, minus maintenance cost, minus AI tool cost, plus the value of defects caught before production. The point of adding terms is not to make the math harder, it's to make the number honest about where automation actually creates or destroys value.
ROI = (Execution_Savings - Maintenance_Cost - AI_Tool_Cost + Defect_Escape_Value) / Total_Automation_Investment
Where:
- Execution_Savings = (Manual_Hours_Per_Cycle - Automated_Hours_Per_Cycle) x Hourly_Rate x Cycles_Per_Period. This is the familiar time-saved term; keep it, it's a legitimate part of the picture, just not the whole picture.
- Maintenance_Cost = Hours spent per period fixing broken tests, updating locators, and triaging flaky failures, x Hourly_Rate. This is the term most ROI calculators skip entirely.
- AI_Tool_Cost = Subscription and usage cost of any AI-assisted test generation, self-healing, or maintenance tooling used in the suite, for the same period. In an AI-assisted testing era, this is a real, recurring line item, not a one-time purchase, and it needs to be netted against the maintenance time it's meant to reduce, not treated as free.
- Defect_Escape_Value = (Defects_Caught_Pre_Production x Average_Cost_Per_Escaped_Defect) minus what that number would have been under the previous (manual, or less mature) process. This is the term that connects ROI to actual quality outcome rather than execution cost alone. Average cost per escaped defect is organization-specific: it should come from your own incident data (support hours, hotfix engineering time, customer-impact estimates), not an industry-wide figure, since publicly cited "cost of a bug" statistics vary enormously by context and are frequently unsourced.
The improvement over time-saved-only ROI is that a suite which is fast to run but expensive to maintain, or one that runs green without actually reducing escaped defects, shows a lower or even negative ROI under this formula, which is the accurate story, not a flattering one.
Coverage percentage tells you what code executed during a test run; it says nothing about whether the assertions in that run would actually fail if the logic were wrong. Mutation testing (deliberately introducing small faults into the code and checking whether the suite catches them) gives a more honest signal of whether coverage is doing real work. A suite with 85% line coverage but a 40% mutation score (`Opinion`, illustrative figures, not a benchmark) is much weaker at catching real regressions than the coverage number alone suggests, and a low mutation score is a leading indicator that your Defect_Escape_Value term is being overstated based on coverage alone. Where a mutation-testing tool is already in your stack (Stryker for JavaScript/TypeScript, PIT for Java, mutmut for Python), pulling a mutation score alongside coverage gives you a more defensible input to the formula than coverage percentage by itself.