Skip to main content
AI Test AutomationQA Best PracticesAI Assisted Testing

How QA Teams Should Actually Use AI in Test Automation

2 September 2026 · OpenCrevo

Search "how to use AI in test automation" and you'll find dozens of near-identical lists: use AI for test generation, use AI for maintenance, use AI for visual regression, and so on. Most of that advice is directionally correct and practically useless, because it never answers the one question that actually determines whether AI helps or quietly hurts your quality bar: who checks the AI's work, and is it a different party than the one who produced it? This article lays out a single organizing principle for AI adoption in QA: use AI freely for speed, generation, and maintenance, but never let the same AI agent that wrote the code (or the tests for that code) be the only thing that decides whether it's correct. That's the self-testing agent problem, and it's the practical pattern most vendor content skips.

The Real Problem

An AI coding agent writes a feature. The same agent, in the same session, with the same context and the same (possibly wrong) understanding of the requirement, then writes the tests that verify that feature. The tests pass. The pull request looks green. Nothing in that loop ever introduced an independent perspective on whether the feature actually does what it's supposed to do.

This is the self-testing agent problem: an AI agent generating both the implementation and its own verification, with no separate check in between. It is structurally similar to a developer writing code and then writing tests that just confirm the code does what it currently does, except AI accelerates the pattern and makes it easier to skip noticing it happened, because the whole loop can complete in minutes and produce a fully green CI run.

The failure mode isn't that AI writes bad tests. It's that AI can write internally consistent, plausible-looking tests and code together, and a test suite generated from the same flawed understanding as the bug will not catch that bug (Industry consensus). If the agent misunderstood the requirement (say, how a discount code should interact with a loyalty tier), it will write a feature that reflects that misunderstanding and tests that confirm the feature works exactly as the agent built it. Passing tests, wrong behavior.

Why This Happens

A few forces push teams toward the self-testing pattern without anyone deciding it as policy:

  • Convenience. The same agent session that just wrote the code is already loaded with the relevant context. Asking it to also write the tests is one more prompt, not a separate workflow step. Nothing forces a pause.
  • Speed pressure. Teams adopting AI in QA are usually trying to move faster. Inserting a separate, independent verification step feels like it reintroduces the bottleneck AI was supposed to remove, so it gets skipped under deadline pressure.
  • A category error about what "AI-assisted" means. Vendor marketing and generic best-practice content frame AI adoption in QA almost entirely around *where* AI helps (generation, maintenance, visual diffing, flaky-test triage) and almost never around *who checks AI's output and how*. That framing gap is a large part of why this specific space is so crowded with listicle-style advice and so thin on the actual governance question (Opinion).
  • No natural forcing function. A human reviewer skimming a PR with 40 new AI-generated test files, all green, has no easy way to tell "verified" tests from "tests that only look like verification." Without a specific check for this, review defaults to trusting the green checkmark.

Common Approaches That Fail

  • "Just review the AI's tests before merging." In principle this works. In practice, a generic code review pass rarely catches whether a test's assertion actually encodes the business rule versus just checking that a response came back. Reviewers are scanning for obvious mistakes, not auditing intent-versus-implementation for every assertion in a 40-file diff.
  • "Use a more capable model." A stronger model reduces some classes of error but does not solve the structural problem: even a highly capable agent still shares the same session context and the same misunderstanding of the requirement across both the code it writes and the tests it writes about that code. Capability isn't independence.
  • "Add more AI-generated tests." More tests generated by the same agent, from the same understanding, cover more of the same ground. Volume doesn't introduce a second perspective; it just multiplies the first one.
  • "Trust coverage percentage as the safety net." Coverage tells you a line executed, not that anyone checked whether the executed behavior was right. A self-testing agent can hit 90+ percent coverage while missing the actual defect, because coverage doesn't ask what the test verified, only what it ran (Industry consensus).
  • Banning AI-generated tests entirely. Some teams respond to this risk by refusing to let AI touch test code at all. This throws away a genuine speed advantage to avoid a governance problem that has a much cheaper fix: keep the generation, add the independent check.

Practical Solution

The organizing principle: separate the agent (or person) that produces work from the agent (or person) that verifies it. AI can sit on either side of that line, but never on both sides of the same piece of work.

Concretely, this means:

  • Generation and verification are different sessions, different prompts, and ideally different actors. If an AI agent wrote the feature and its tests in one session, the verification pass must come from somewhere else: a different agent with a fresh context and no visibility into the first agent's assumptions, or a human reviewer, or (for anything customer-facing or compliance-relevant) both.
  • The verification pass has a specific job, not a generic "review this" prompt. It should be asked directly: does this test verify the requirement, or does it verify what the code currently does? Would this test fail if the underlying business rule were subtly wrong, not just if the code crashed?
  • Independence is a property of the process, not a one-time audit. This isn't a single pre-launch review. It's a repeatable step that runs on every PR that includes AI-generated code, AI-generated tests, or both.
  • Where the stakes are highest (customer-facing logic, anything with a compliance or financial consequence), the independent check should be human-led, not just a second AI pass, because a second AI agent can still share systemic blind spots with the first if both were trained or prompted the same way (Opinion).

A useful gut check for any AI-in-QA workflow: if you removed the AI entirely and described the process to a compliance auditor, would "the same person wrote the code and signed off on its own tests" pass as a control? If the answer is no for a human, it's no for an AI agent standing in for that human, and the workflow needs a second, independent party in the loop.

Implementation

A concrete pattern for a Playwright/TypeScript project using an AI coding agent (Cursor, Claude Code, GitHub Copilot, or similar):

  • Agent A generates the feature and an initial test draft, in its normal working session.
  • The PR is explicitly labeled (a simple GitHub label like `ai-generated` or `needs-independent-verification`) so the verification step isn't left to memory. A label-based gate is cheap to implement and easy for a reviewer or CI check to enforce.
  • A separate verification pass runs against the labeled PR, using either: a fresh AI agent session with no access to Agent A's chat history or reasoning, prompted specifically to check whether each test asserts the business rule or just the current implementation, or a human reviewer using a short, specific checklist (see below) rather than a general "LGTM" pass.
  • CI enforcement, so this isn't optional: a GitHub Actions check that blocks merge on PRs carrying the `ai-generated` label until a second reviewer (human or a distinctly configured verification agent) has approved, separate from the author's own approval.

A minimal GitHub Actions snippet enforcing a second, distinct approval on labeled PRs:

name: independent-verification-gate
on:
 pull_request:
 types: [labeled, synchronize, opened]

jobs:
 require-second-reviewer:
 if: contains(github.event.pull_request.labels.*.name, 'ai-generated')
 runs-on: ubuntu-latest
 steps:
      - name: Check for independent approval
 uses: actions/github-script@v7
 with:
 script: |
 const reviews = await github.rest.pulls.listReviews({
 owner: context.repo.owner,
 repo: context.repo.repo,
 pull_number: context.payload.pull_request.number,
            });
 const approvals = reviews.data.filter(r => r.state === 'APPROVED');
 const author = context.payload.pull_request.user.login;
 const independentApproval = approvals.find(r => r.user.login !== author);
 if (!independentApproval) {
 core.setFailed('PR is labeled ai-generated and requires an approval from someone other than the author.');
            }

This is a starting point, not a complete governance system: it enforces that *someone* other than the author approved, not that the approver actually ran the specific verification checklist. Pair it with the checklist below as a PR template section, so the reviewer has a concrete job rather than a vague sign-off.

A short reviewer checklist to paste into the PR template for anything labeled `ai-generated`:

  • Does each new assertion check the business rule, or just that a response came back / no error was thrown?
  • If the underlying logic were subtly wrong (not crashed, just wrong), would this test fail?
  • Was this test written by the same session that wrote the code it tests? If yes, has an independent pass happened yet?
  • For anything customer-facing or compliance-relevant, has a human (not just a second AI pass) signed off?

AI Considerations

AI is genuinely good at the first half of this pattern: fast generation, boilerplate test scaffolding, maintaining existing suites as UI or API shapes shift, and flagging likely flaky tests for triage. None of that needs to slow down. The mistake isn't using AI for those things; it's collapsing generation and verification into the same actor and the same pass.

It's also worth being specific about a subtlety: using a second AI agent as the independent check is a real improvement over no check at all, but it is not equivalent to human review for anything high-stakes. Two AI agents, especially if built on similar models or prompted with similar instructions, can share the same blind spots (Opinion, requires verification: there is no published study quantifying how often two independent LLM agents converge on the same wrong assumption for a given codebase, so treat this as a risk to manage conservatively rather than a measured probability). Practically: use a second AI pass as a first-line, always-on check across all PRs, and reserve human-led review specifically for the requirement categories where a shared blind spot would be expensive: payment logic, compliance-relevant flows, anything with a safety implication, and anything a self-testing agent has been wrong about before.

OpenEvident

If you want to see the "agent writes code, local tooling drives the actual test run" half of this pattern in a real, self-hostable project rather than building it from scratch, OpenEvident's CrevoAI project is a local-first Playwright test automation toolkit built specifically for AI-coding-agent workflows (Cursor, Claude Code, GitHub Copilot). Its architecture keeps codegen, browser control, and test execution on your own machine via a local MCP server, with no cloud job runner or remote orchestration involved. That local-first design is relevant here for a narrower reason than "it generates tests": it keeps the artifacts a verification pass needs to inspect (the actual browser recordings, the actual generated test files) on infrastructure your team controls, rather than routed through a third-party service, which matters if audit trail and data residency are part of how you'd implement the independent-verification pattern above. `[VERIFY WITH OPENEVIDENT TEAM]`: whether CrevoAI's own MCP tools include any built-in flagging for tests written by the same session as the code under test, versus this being a workflow a team layers on top of it, as described in this article.

OpenCrevo Implementation

The core idea in this article, that AI-generated work needs a genuinely separate verifier and not a self-check, is the exact premise behind OpenCrevo's Independent Testing service: independent, human-led verification against AI-generated code and outputs, catching what self-testing agents miss. If your team has already adopted AI for test generation and maintenance and now needs to decide how to structure the second, independent pass described above (rather than bolting on ad hoc review that quietly erodes under deadline pressure), OpenCrevo builds that verification layer into an existing pipeline without replacing the AI tooling you're already using for generation and maintenance. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Never let the same AI session (or the same person) that wrote a feature be the sole approver of the tests written for that feature.
  • Label PRs containing AI-generated code or tests so the verification step is enforced, not left to memory.
  • Give the verification pass a specific job: does this assertion check the business rule, or just that something ran?
  • Enforce independent approval in CI for labeled PRs, not just as a review-culture expectation.
  • Reserve human-led (not just second-AI) verification for customer-facing, compliance-relevant, or previously-wrong-before categories of change.
  • Keep using AI freely for generation, maintenance, and triage; the fix is adding a separate check, not slowing down generation itself.

FAQ

  • Does this mean every AI-generated test needs a human reviewer? No. A second, independent AI pass with a specific verification prompt is a reasonable first-line check for most PRs. Reserve human-led review for higher-stakes categories: customer-facing logic, compliance-relevant flows, and anything a self-testing agent has gotten wrong before.
  • Is this worth it for a small team? The label-plus-CI-gate pattern above is lightweight enough for a small team to run without hiring anyone new; it mainly costs a PR template addition and a short GitHub Actions job. The organizational discipline (never letting one session both write and self-certify) scales down fine; what doesn't scale down well is skipping it under deadline pressure, which is exactly when it matters most.
  • How is this different from just "reviewing AI-generated tests" as general advice? General review advice doesn't specify who does the review or what they're checking for. This pattern is specific: the reviewer must be a genuinely separate party (not the same session/agent that produced the work) and must check a specific thing (does the assertion verify intent, not just current behavior).
  • Does using a second AI agent as the check fully solve the problem? It's a real improvement over no check, but not equivalent to human review for high-stakes changes, because two AI agents can share the same blind spots (Opinion). Use it as an always-on first-line check and add human review for the categories that would be expensive to get wrong.
  • What's the difference between this and mutation testing? They're complementary, not competing. Mutation testing is a technical check on whether a test would catch a subtly wrong implementation. The self-testing agent problem is a process/governance issue: who is allowed to sign off on whose work. A mutation-testing pass run by the same agent that wrote the code is still a self-check; run by an independent party, it becomes part of the verification pattern this article describes.

Conclusion

The generic "how to use AI in QA" advice isn't wrong so much as incomplete: it tells you where AI can help and skips the one question that determines whether that help is safe, which is whether anything independent ever checks the AI's own output. Use AI for speed everywhere it earns it: generation, maintenance, triage. Just never let the agent that produced the work also be the only thing that grades it. That single separation, enforced as a repeatable process rather than a one-time review, is the practical difference between AI adoption that holds up under scrutiny and AI adoption that looks fast until the first missed defect reaches production.

For the two workflow questions this article assumes but doesn't fully answer, see 5 Things You Must Validate Before Trusting an AI-Generated Test for the specific validation checklist referenced above, and How to Introduce AI Into QA Without Turning Your Test Suite Into a Black Box for the broader governance and audit-trail framework this pattern fits inside. If you're planning this at the level of an entire team's transition rather than a single workflow, From Manual Regression to AI-Assisted Quality Engineering: A Practical Migration Plan covers the org-level rollout sequence.

Sources

OpenEvident's vindicate repository, OpenCrevo services documentation (src/data/services.ts in this repository), industry discussion of coverage-versus-verification gaps in AI-generated test suites (Industry consensus, synthesized from current QA-engineering practitioner commentary, not a single primary source), GitHub Actions documentation for pull request review APIs.

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.