- Explore with @playwright/cli (or a short MCP session you discard).
- Commit an eval fixture with gold chunk IDs — see RAG evaluation metrics.
- Never let the explore agent be the merge gate.
# AI testing explore — local only
npx @playwright/cli@latest open https://staging.example.com/support
npx @playwright/cli@latest snapshot
# AI testing gate — CI
npx playwright test tests/eval.spec.ts --project=eval
import { test, expect } from "@playwright/test";
import { complete, retrieve } from "../src/rag";
test("refund answer cites policy-refunds-v3", async () => {
const q = "What is the refund window?";
const ids = await retrieve(q);
const out = await complete(q);
expect(ids).toContain("policy-refunds-v3");
expect(out.toLowerCase()).toContain("14 days");
expect(out.toLowerCase()).not.toContain("30 days");
});
Productionise the gate through AI quality engineering services or the Enterprise page. Score your current baseline first with the free QA maturity assessment.