The Real Problem
A developer merges an AI-generated Playwright test. It passes in CI. Three weeks later, the feature it covers breaks in production, and the test is still green, because it was never actually testing the thing that broke. Nobody flagged the test as AI-generated during review, so nobody applied any extra scrutiny to it, and it looked exactly like every other test in the file.
This is not a hypothetical. It is the predictable outcome of treating a generated test the same way you'd treat a hand-written one: reviewed for syntax and whether it passes, not for whether it actually verifies the right thing. The fix is not "review AI-generated tests more," which doesn't scale and burns out reviewers. The fix is a short, specific, repeatable set of checks applied to every AI-generated test, every time, before it merges.