Skip to main content
AI Test ValidationTest Review ChecklistQA Sign-Off Process

5 Things You Must Validate Before Trusting an AI-Generated Test

2 September 2026 · OpenCrevo

An AI agent just wrote you a test. It's syntactically correct, it runs, and it's green. None of that tells you whether it should be merged. This article is the fast, practical companion to the deeper question of why AI-generated tests create false confidence (/blog/ai-generated-tests-false-confidence/): instead of the argument, it's the checklist. Five specific things to check before an AI-generated test enters your suite: assertion strength, negative-case coverage, repeat-run stability, alignment with the actual requirement, and a governance step most teams skip entirely, a PR label and sign-off trail that makes AI-generated tests visible and accountable in review. None of these checks take more than a few minutes each. Skipping all five is how a green suite ships a real bug.

The Real Problem

A developer merges an AI-generated Playwright test. It passes in CI. Three weeks later, the feature it covers breaks in production, and the test is still green, because it was never actually testing the thing that broke. Nobody flagged the test as AI-generated during review, so nobody applied any extra scrutiny to it, and it looked exactly like every other test in the file.

This is not a hypothetical. It is the predictable outcome of treating a generated test the same way you'd treat a hand-written one: reviewed for syntax and whether it passes, not for whether it actually verifies the right thing. The fix is not "review AI-generated tests more," which doesn't scale and burns out reviewers. The fix is a short, specific, repeatable set of checks applied to every AI-generated test, every time, before it merges.

Why This Happens

Three things compound here. First, AI-generated tests look identical to hand-written ones once they're in the codebase: same syntax, same file, same CI job. There's no visual signal to a reviewer that this test deserves a different level of scrutiny `(Opinion)`, unless the team deliberately creates one. Second, generated tests are optimized to pass against the current implementation, so a test that would fail against a subtly wrong implementation and a test that only proves the code ran look the same at a glance. Third, most existing "review your AI tests" guidance (Katalon, testRigor, and similar vendor content) stops at generic advice like "read the test before merging it," without a concrete, five-minute checklist a reviewer can actually run, and without addressing the governance question of who is accountable for having done that review at all.

Common Approaches That Fail

  • "Read it before merging." True but not actionable. Read it for what, specifically? Without a defined checklist, "review" becomes whatever the reviewer happens to notice that day.
  • Treating a passing CI run as validation. A test passing tells you the test ran without an error. It does not tell you the test would fail if the underlying logic were wrong.
  • Applying the same review bar to every test regardless of origin. AI-generated tests fail in a specific, different pattern (happy-path bias, weak assertions, no negative cases) than hand-written ones. A generic review checklist misses the failure mode that's actually common here.
  • No record of which tests were AI-generated. Six months later, when a bug slips through, nobody can even identify whether the responsible test was AI-written or hand-written, so the team can't learn from the pattern or tighten the process.

Practical Solution

Run these five checks on every AI-generated test before it merges. None require special tooling beyond what a typical Playwright/CI setup already has.

  • Assertion strength. Open the test and read every assertion, not just the test name. `expect(response.status).toBe(200)` or `expect(result).toBeDefined()` confirms the code ran, not that it did the right thing. A test asserting on a checkout total should assert the actual calculated value, not just that a value exists. If you can't describe, in one sentence, what business rule this assertion would catch a violation of, it's too weak to trust `(Opinion)`.
  • Negative-case coverage. Check whether the AI generated any test for the input that shouldn't work: the invalid input, the conflicting state, the unauthorized user. AI-generated suites are structurally biased toward the happy path `(Industry consensus)`, because the model is reasoning from what the code does, not from what a user might do wrong. If every generated test in a batch is a success-path test, that's a signal, not a coincidence.
  • Flakiness on repeat runs. Run the new test 10 to 20 times in a loop (`npx playwright test --repeat-each=20` for a single spec file) before merging it. A test that passes once but fails intermittently under repetition is often relying on a race condition or a timing assumption the AI agent happened to get away with during generation. Catching this before merge is far cheaper than debugging a flaky test that's already been in the suite for months.
  • Alignment with the actual acceptance criteria, not just the code. This is the check most teams skip because it requires looking away from the code entirely. Pull up the actual requirement, ticket, or acceptance criteria the feature was built against, and ask: does this test verify that criteria, or does it verify whatever the code currently happens to do? An AI agent reading only the source file has no access to intent it wasn't told; it can only test what it can observe. If the requirement says "a user cannot apply two discount codes to the same order" and no generated test attempts that, the test suite is aligned with the code, not with the requirement, and that gap is exactly where production bugs live.
  • PR labeling and audit-trail sign-off. This is the governance step existing review-checklist content (Katalon, testRigor, and similar) doesn't cover, and it matters most at enterprise scale, where dozens of AI-generated tests can merge per week across multiple teams. Concretely: Apply a required label, e.g. `ai-generated-test`, to any PR that includes AI-generated test code. Some teams enforce this with a lightweight CI check that greps the diff for a generation-tool marker or requires the author to self-tag it; either way, the label needs to be non-optional, not left to memory. Require a named human reviewer sign-off specifically against this checklist, separate from a general code-review approval. A `CODEOWNERS` entry on test directories, or a required review group in your PR settings, works for this. Log which tests were AI-generated and who signed off, even if it's just a line in the PR description or a tag in your test-management tool. When a bug ships despite a passing suite, this is what lets you trace back to whether the responsible test was reviewed against the checklist above, or waved through.

Implementation

A minimal version of checks 3 and 5, wired into GitHub Actions:

name: ai-generated-test-gate
on:
 pull_request:
 paths:
      - "tests/**"

jobs:
 repeat-run-check:
 if: contains(github.event.pull_request.labels.*.name, 'ai-generated-test')
 runs-on: ubuntu-latest
 steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npx playwright install --with-deps
      - name: Run new/changed spec 15 times to catch flakiness
 run: npx playwright test tests/checkout.spec.ts --repeat-each=15

 require-signoff:
 if: contains(github.event.pull_request.labels.*.name, 'ai-generated-test')
 runs-on: ubuntu-latest
 steps:
      - name: Block merge without designated reviewer approval
 run: echo "Enforced via branch protection: require review from @qa-leads on ai-generated-test PRs"

The `repeat-run-check` job only runs when the `ai-generated-test` label is present, so it doesn't slow down every PR, only the ones that need the extra scrutiny. The sign-off requirement itself is enforced through branch protection rules and a `CODEOWNERS` entry, not through the workflow file, since GitHub Actions can't itself force a specific reviewer, `[VERIFY API BEFORE PUBLICATION]` for the exact branch-protection configuration syntax in your own repo settings.

AI Considerations

The same agent that generated the test can usually be prompted to fix what these checks surface, once the gap is stated concretely. "This test only checks `toBeDefined()`, rewrite it to assert the actual expected discount value" or "write a test for the case where two discount codes are applied to the same order" are specific, actionable prompts an agent can act on directly. The checklist doesn't replace the AI agent's usefulness, it turns a vague "review this" into instructions the agent can execute, which keeps a human in the loop without turning every review into a manual rewrite.

OpenEvident

If your team is generating Playwright tests through an AI coding agent (Cursor, Claude Code, GitHub Copilot) in the first place, the checklist above applies most directly to that workflow. OpenEvident's CrevoAI project is a local-first Playwright automation toolkit built specifically for AI-agent-driven test generation: a VS Code/Cursor extension talks to a local MCP server that drives codegen and browser control, with no cloud job runner and no remote orchestration involved. That local-only architecture matters for check 5 above specifically, since audit-trail and sign-off requirements are easier to enforce when the generation step itself isn't sending your test data through a third-party service. `[VERIFY WITH OPENEVIDENT TEAM]`: whether CrevoAI exposes any built-in tagging of AI-generated output for the kind of PR-label workflow described here, versus this being a step a team layers on top of it.

OpenCrevo Implementation

Running this checklist consistently across every team, every PR, every week is where most organizations lose the thread, not because the checks are hard, but because nobody owns enforcing them once the initial excitement about AI test generation wears off. OpenCrevo's Independent Testing service is built around exactly this gap: independent, human-led verification against AI-generated code and outputs, catching what self-testing agents miss. If your team needs this checklist applied by someone other than the same agent (or the same team) that generated the tests in the first place, OpenCrevo can build that independent verification layer into your existing pipeline. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Assertion strength: every assertion checks an actual expected value or business rule, not just "no error" or "status 200."
  • Negative-case coverage: at least one generated test per feature covers an invalid input, conflicting state, or unauthorized action, not only the happy path.
  • Repeat-run stability: new test run 10 to 20 times (`--repeat-each`) before merge, with zero intermittent failures.
  • Requirement alignment: test verified against the actual acceptance criteria or ticket, not just against what the code currently does.
  • PR labeling and sign-off: `ai-generated-test` label applied, named reviewer approval required, and the decision logged somewhere retrievable later.

FAQ

  • Does every AI-generated test need all 5 checks? For anything touching a critical path (checkout, auth, payments, data mutation), yes. For low-risk UI cosmetics, teams can reasonably apply a lighter version, `(Opinion)`, but the label-and-sign-off step is cheap enough to apply universally regardless of risk tier.
  • How long does this add to review time? A few minutes per test once the checklist is habitual; the repeat-run check adds CI time (roughly 10 to 20x a single test's normal runtime) but only for labeled PRs, not the whole suite.
  • Does this replace mutation testing? No, they solve different problems. Mutation testing (covered in the false-confidence article above) tells you systematically whether your suite would catch a wrong implementation. This checklist is the fast, per-PR gate that catches the same problem and adds the governance layer mutation testing doesn't provide on its own.
  • What if our team doesn't use GitHub? The mechanism (a required label, a scoped repeat-run job, a required reviewer group) maps onto GitLab merge-request labels and approval rules, or Bitbucket's equivalent, with the same underlying logic.
  • Who should be the required reviewer for AI-generated tests? A QA lead or senior SDET familiar with the checklist, not necessarily the original feature developer, since the point of the sign-off is an independent second look, not a rubber stamp from whoever wrote the code the test covers.

Conclusion

An AI-generated test earns trust the same way any test does: by demonstrating it would actually fail if the underlying logic were wrong. The difference is that AI-generated tests fail in a specific, predictable pattern (weak assertions, missing negative cases, happy-path bias) that a five-item checklist can catch in minutes, provided someone is actually required to run it. The label-and-sign-off step is what turns "we review AI tests" from an intention into an enforced, auditable practice. For the deeper argument behind why this matters, see Why AI-Generated Tests Create False Confidence; for how generation quality varies by tool and prompt, see AI Test Generation Is Easy, Reliable Test Generation Is Not; for the broader adoption question, see How QA Teams Should Actually Use AI in Test Automation.

Sources

Sources include Playwright documentation on --repeat-each and test retries, GitHub Actions documentation on pull_request label conditions and branch protection, industry discussion of AI-generated test happy-path bias (Industry consensus, synthesized from multiple current QA-engineering publications, not a single primary source), OpenEvident's vindicate repository, and OpenCrevo services documentation (src/data/services.ts in this repository).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.