Skip to main content
AI Test GenerationTest ReliabilityLLM Test Quality

AI Test Generation Is Easy, Reliable Test Generation Is Not

2 September 2026 · OpenCrevo

Ask any AI coding agent to "write tests for this function" and it will comply in seconds, producing a file full of green checkmarks. That part is genuinely solved. What is not solved, and what most teams discover only after a defect ships, is that the ease of generating a test has almost no relationship to the reliability of that test. Reliability isn't something you bolt on after the fact by running a coverage report or a mutation-testing pass, though those checks matter (see the companion article on validating AI-generated tests (/blog/validate-ai-generated-test-checklist/) for that side of the problem). Reliability is mostly decided at the moment the test is generated: what the agent was given to work from, how the prompt was structured, and whether anyone asked it to think about failure before it thought about success. This article is about that moment, the generation workflow itself, not the audit that happens afterward.

The Real Problem

A team adopts an AI coding agent to speed up test writing. Within a sprint, test file counts double. Everyone is happy. Then a request comes in: "why didn't we catch the refund-after-partial-shipment bug?" The answer, when someone actually goes and reads the generated tests, is unremarkable and completely predictable: every generated test fed the function a normal, in-range input and asserted a normal, in-range output. Nobody, human or AI, ever asked the generation step to produce a test for what happens when the shipment is only half-fulfilled and a refund request arrives mid-flight.

This is not a coverage gap that a coverage report would have shown you (Industry consensus). Coverage tools measure whether a line executed, not whether anyone asked the interesting question about that line. The refund function was fully covered. It was also fully untested against the one input that mattered. The generation process produced a test that was easy to write and cheap to pass, and reliability was never actually part of the brief.

Why This Happens

Two separate mechanics are at play here, and conflating them is why "just use a better model" doesn't fix the problem on its own.

  • The agent generates from what it can see, not from what the feature is supposed to do. An AI agent asked to write tests for a function typically reads the function's implementation, not the ticket, the acceptance criteria, or the product requirement that motivated the function in the first place. A test generated purely from source code can only ever describe what the code currently does. It has no independent source of truth to check that behavior against, so it cannot catch the case where the code itself is wrong. This is a structural limitation of code-only prompting, not a model-quality problem (Industry consensus).
  • Nobody explicitly asked for the negative case. Left to its own judgment, a generation prompt like "write tests for this endpoint" produces the tests that require the least reasoning: valid input in, expected output out. Generating a meaningful edge case (a boundary value, a malformed payload, a race condition, a permission the caller shouldn't have) requires the agent to reason about intent and failure modes that were never stated. Models are capable of this reasoning (Opinion), but only when the prompt or workflow actually asks for it. An unscoped, single-shot generation request rarely does.

The common thread: reliable test generation is a requirements problem and a prompting problem before it is ever a validation problem. Article 01 in this series covers what happens when you skip validation and how a mutation-testing pass exposes the resulting gap after the fact. This article is about not creating that gap in the first place.

Common Approaches That Fail

  • "Point the agent at the code and let it write tests." This is the single most common generation pattern and the single most reliable way to reproduce happy-path bias, because the code is the only source of truth the agent has, and code cannot tell you when it's wrong about its own intent.
  • One giant prompt asking for "comprehensive test coverage." Vague instructions produce vague results. An agent told to be "comprehensive" without a concrete list of scenario categories will default to the scenarios it can generate with the least effort: the same happy-path shapes, repeated with different input values.
  • Generating everything in one pass and reviewing afterward. By the time a human reviewer sees 40 generated tests, reviewing them thoroughly enough to notice a systematic gap (not "this test has a bug" but "this whole batch never tests concurrent access") is a much harder review task than shaping the generation request correctly the first time.
  • Assuming a newer or larger model fixes this. A more capable model still needs to be told what "correct" means for this specific feature and still needs to be asked, explicitly, to consider failure. Model capability changes how well it executes a well-scoped request; it does not supply the scope on its own (Opinion).

Practical Solution

The fix lives upstream of validation, in how the generation request itself is built. Three concrete changes to the generation workflow, in order of impact:

1. Feed the agent the requirement, not just the code. Before asking for tests, give the agent the actual acceptance criteria, user story, or written requirement, alongside the implementation, not instead of it. A prompt structured as "here is the acceptance criteria for this feature: [AC text]. Here is the implementation: [code]. Generate tests that verify the acceptance criteria is met, not just that the code runs" gives the agent an independent source of truth to check the code against. Without that second source, the agent can only ever validate the code against itself, which is not validation at all, it's restatement.

2. Generate negative and edge cases as an explicit, separate step, with a named category list. Instead of one open-ended request, split generation into passes with a fixed checklist of scenario categories the agent must address for each function or endpoint under test:

  • Boundary values (empty, zero, maximum, one past maximum)
  • Invalid or malformed input (wrong type, missing required field, malformed format)
  • State-dependent failure (the operation is attempted in a state it wasn't designed for, like a refund on an order that already shipped)
  • Concurrent or out-of-order operations (two requests that touch the same resource at nearly the same time)
  • Authorization boundaries (a caller who shouldn't be allowed to perform this action attempts it anyway)

Asking "does this feature have a boundary-value test? A malformed-input test? A state-conflict test?" as five separate, named questions produces noticeably more of each category than one instruction to "be thorough" (Industry consensus, consistent with how structured prompting generally outperforms open-ended prompting for enumeration-style tasks).

3. Generate one test per stated risk, not N tests for volume. Before generation, write down (or have the agent draft, then a human confirm) the specific things that could go wrong with this feature in production terms, not test-framework terms. "Two discount codes applied in the same session" is a risk statement. "More edge case tests" is not. Each generation request should map to one risk statement, which keeps the output auditable: a reviewer can check the risk list against the generated tests directly, rather than reading 40 tests trying to reverse-engineer what they were meant to prove.

Implementation

A concrete before-and-after for a Playwright test generation prompt, for a refund endpoint:

Unreliable generation prompt (code-only, no requirement, no scenario list):

Write Playwright API tests for the /refunds endpoint based on this handler code: [code].

Reliable generation prompt (requirement included, scenario categories named explicitly):

Acceptance criteria (from ticket REFUND-114):
- A refund can only be issued against an order in "delivered" or "partially_shipped" status.
- A refund amount cannot exceed the sum of items actually shipped.
- A refund request against an order still "in_fulfillment" must be rejected with a 409, not silently queued.

Implementation: [code]

Generate Playwright API tests for /refunds covering each of the following categories explicitly,
one test per named scenario, and reference which acceptance criteria line each test verifies:
1. Boundary: refund amount exactly equal to shipped total.
2. Invalid state: refund attempted while order is "in_fulfillment".
3. Partial shipment: refund attempted for more than the shipped subset.
4. Authorization: refund attempted by a user who does not own the order.

The second prompt produces tests that trace back to a specific acceptance-criteria line, which is a workflow benefit beyond just better coverage: it gives a reviewer (or an audit trail) a direct way to check "does a test exist for AC-2" without reading test internals line by line. That traceability discipline is the same practice the story-driven test generation described in the 5 things you must validate before trusting an AI-generated test checklist depends on, it's just applied one step earlier, at generation time instead of review time.

A minimal Playwright test generated from category 2 above:

test('rejects refund when order is still in_fulfillment [AC-3]', async ({ request }) => {
 const order = await createOrder({ status: 'in_fulfillment', shippedTotal: 0 });

 const response = await request.post(`/refunds`, {
 data: { orderId: order.id, amount: 50 },
  });

 expect(response.status()).toBe(409);
 const body = await response.json();
 expect(body.error).toContain('order not eligible for refund');
});

Note the assertion checks the specific rejection reason, not just the status code, because a test that only checks `response.status()).toBe(409)` would still pass if the endpoint rejected the request for the wrong reason entirely, which is exactly the kind of shallow assertion this generation discipline is meant to prevent.

AI Considerations

The same agent capability that makes generation fast can be redirected toward reliability, but only if the workflow asks it to. A useful pattern: after the agent generates a first pass of tests, run a second, adversarial generation pass with a prompt like "review the tests you just generated. For each one, identify what change to the implementation's business logic (not a typo) would NOT be caught by this test. Then generate one additional test to close that gap." This turns the agent into its own first-line reviewer, catching the shallow-assertion problem before a human or a mutation-testing tool has to catch it later. It is a generation-time practice, not a validation-time one, because it happens as part of producing the test suite, not as a separate audit pass afterward. It also does not replace mutation testing, it just means fewer mutants survive by the time that check runs.

OpenEvident

CrevoAI, the Playwright test automation toolkit published under OpenEvident, is built specifically to run inside an AI coding agent's workflow (Cursor, Claude Code, GitHub Copilot, and similar tools), with MCP tools for codegen, browser control, and recordings, entirely on the developer's own machine. Because the generation loop happens locally, with no cloud job runner in between, a team can iterate on the kind of scenario-driven, requirement-fed prompting described above directly inside the same tool the agent already uses, rather than round-tripping generated code through a separate service. `[VERIFY WITH OPENEVIDENT TEAM]`: whether CrevoAI's own codegen tooling exposes a structured way to attach acceptance criteria or a named scenario checklist to a generation request, versus this being a prompting discipline a team layers on top of it themselves.

OpenCrevo Implementation

Fixing generation-time practices is a workflow and standards problem as much as a tooling one: someone has to define the scenario categories, decide what "acceptance criteria" looks like in a form an agent can consume, and make sure the discipline survives beyond the one engineer who currently does it well. OpenCrevo's Independent Testing service, built around independent, human-led verification against AI-generated code and outputs, exists to catch exactly what a team's own generation habits miss, and to help formalize the requirement-fed, scenario-driven generation pattern described here into something the whole engineering organization actually follows. If you'd like a second, unbiased look at how your team's AI-generated tests are actually produced, not just how they're reviewed afterward, OpenCrevo can help build that into your existing pipeline. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Before generating a test, attach the actual requirement or acceptance criteria to the prompt, not just the implementation code.
  • Split generation into named scenario categories (boundary, invalid input, state conflict, concurrency, authorization) instead of one open-ended "write tests" request.
  • Map each generated test back to a specific risk statement or acceptance-criteria line, not a vague "more coverage" goal.
  • Run a second, adversarial generation pass asking the agent what change to the business logic its own tests would miss.
  • Check that assertions verify the specific reason for a result (an error message, a business rule), not only a status code or "no exception thrown."
  • Treat a test with no traceable acceptance-criteria link as a candidate for review or deletion, not as free coverage.

FAQ

  • Is AI test generation reliable out of the box? Not by default (Industry consensus). Reliability depends on what the agent is given at generation time (the requirement, not just the code) and whether the prompt explicitly asks for negative and edge cases. An unscoped "write tests for this" request reliably produces happy-path-biased output regardless of model quality.
  • Does giving the agent acceptance criteria replace mutation testing? No, they solve different problems. Feeding acceptance criteria at generation time reduces how many gaps exist in the first place. Mutation testing, covered in the companion article on false confidence from AI-generated tests, catches what still slips through afterward. Use both.
  • How long does it take to add a scenario checklist to an existing generation workflow? Writing the initial category list (boundary, invalid input, state conflict, concurrency, authorization, or a set tailored to your domain) is typically a single working session. Getting the team to consistently attach requirements to generation prompts is more of a habit change than a tooling change, and takes longer to stick.
  • What breaks first when a team tries this at scale? Usually the requirement side, not the prompting side (Opinion): many teams don't have acceptance criteria written down in a form that's easy to hand to an agent, so the first real blocker is a documentation gap, not an AI capability gap.
  • Does this replace human test design entirely? No. It changes what a human needs to supply (a clear requirement and a named list of risk categories) rather than removing the human from the loop. The autonomous QA article in this series covers how far that loop can actually be automated today and where it currently breaks.

Conclusion

Generating a test that passes is trivial. Generating a test that would have caught the bug your team actually shipped is a different task, and it's decided upstream of any review or mutation-testing step, at the moment the generation request is written. Feed the agent the requirement, not just the code. Ask for negative and edge cases by name, not by vibe. Map each test to a stated risk instead of a raw count. Do that, and the validation work covered elsewhere in this series has far less to catch, because reliability was built in at generation time instead of chased down after the fact.

Sources

Industry discussion of happy-path bias and structured-versus-open-ended prompting for test generation (Industry consensus, synthesized from multiple current QA-engineering and prompt-engineering publications, not a single primary source), OpenEvident's vindicate repository, OpenCrevo services documentation (src/data/services.ts in this repository), Playwright documentation.

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.