Skip to main content
Playwright Locator StrategyBrittle CSS SelectorsGetByRole Vs Data-Testid

Stop Using Fragile CSS Selectors: A Practical Locator Strategy for Modern Web Apps

2 September 2026 · OpenCrevo

Every Playwright tutorial tells you the same thing: stop using CSS class selectors, use `getByRole`, fall back to `data-testid`. That advice is correct and also incomplete, because it treats every application as if it has the same starting conditions. It doesn't. A greenfield React app built with shadcn/ui and proper ARIA semantics is a completely different locator problem than a ten-year-old jQuery admin panel with div soup and no accessibility tree to speak of. This article gives you a decision framework for choosing a locator strategy based on what your application actually gives you to work with (an accessibility audit, effectively), what it costs to retrofit test IDs into legacy markup, and how each locator type actually behaves under real CI conditions, not just in a demo repo. The goal is a strategy you can apply to the specific app in front of you, not a rule you copy from a blog post.

The Real Problem

A team migrates its Playwright suite off `page.locator('.btn-primary.submit-form')` and onto `page.getByRole('button', { name: 'Submit' })`, because every guide says role-based locators are the resilient choice. Three sprints later, half the new locators are still breaking on every deploy, just with a different error message. The reason: the app's buttons don't have consistent accessible names, several interactive elements are unlabeled `<div onClick>` handlers with no role at all, and the design system swaps visible button text based on A/B test flags. `getByRole` didn't fail because the advice was wrong. It failed because the team applied a locator strategy the app's own markup couldn't support, and nobody checked that first.

This is the actual gap in most locator-strategy content: it presents `getByRole` as a universal best-first choice, when in practice its resilience is entirely conditional on the application's accessibility semantics being reasonably correct. An app with a real accessibility tree makes `getByRole` close to free. An app without one makes `getByRole` a locator that looks resilient in the Playwright docs and is actually just as brittle as a CSS class selector, because you are still locating by presentation-layer text and structure, just through a different API.

Why This Happens

Three separate forces converge to produce this problem:

  • Locator guides are written against demo apps, not your app. Playwright's own documentation and most third-party guides (Industry consensus: widely cited by bug0, momentic, and similar comparison posts) use clean example markup with correct roles and labels already in place. That's a reasonable teaching example. It's a poor proxy for a real production app that accumulated five years of ad hoc component patterns before anyone thought about test automation.
  • Accessibility and testability are built by different incentives, at different times. A component library gets ARIA roles added because a compliance requirement or a screen-reader bug report forced it, not because a test author asked for it. Many production apps have partial, inconsistent accessibility coverage: some flows are fully labeled, others (often the newest or most custom-built ones) have none. A locator strategy that assumes uniform accessibility semantics across the whole app will work in the well-labeled 60 percent and quietly fail in the rest.
  • Retrofitting is treated as free when it isn't. Adding `data-testid` attributes to existing markup sounds like a five-minute change per element. At scale, it means touching component source across every team that owns a screen under test, getting those PRs reviewed and merged, and doing it without introducing a naming collision or duplicate ID (`Opinion`, based on how this typically plays out in a codebase with more than one contributing team). For a legacy app where the original authors are gone and the component library isn't even the one currently favored by the team maintaining it, this cost is often the real reason locator migrations stall, not a lack of will.

Common Approaches That Fail

  • "Just switch everything to getByRole." Works cleanly on apps with a genuinely correct accessibility tree. On apps without one, it just moves the brittleness from a CSS class name to an accessible name that's equally likely to change, while giving the team false confidence that they've "done the resilient thing."
  • "Just add data-testid everywhere." Solves the resilience problem but ignores that `data-testid` is invisible to assistive technology and to real users, so it tells you nothing about whether the feature is actually accessible. It also requires source-code changes to every component under test, which is a real engineering cost, not a QA-only task, and gets deprioritized against feature work in exactly the codebases that need it most.
  • "Use XPath, it can find anything." True, and that's the problem. An XPath expression that walks the DOM by structure (`//div[3]/span[2]/button`) is the single most fragile locator category that exists: any structural change in an unrelated part of the layout, a new wrapper div, a reordered sibling, breaks it, even when the element itself hasn't changed at all. XPath's real place in a locator strategy is much narrower than "use it when the others fail."
  • Treating locator choice as a one-time decision for the whole app. Real applications are not uniform. Treating "our locator strategy" as a single answer for the entire codebase, rather than a per-flow or per-component decision informed by what that specific part of the app actually supports, is the root cause of most of the failures above.

Practical Solution

The original contribution of this article is this: locator strategy should follow a decision order driven by what your accessibility posture already tells you about a given screen or component, not by a fixed universal ranking applied uniformly everywhere.

The decision framework, per component or flow, in priority order:

  • Does this element have a correct, stable ARIA role and accessible name already, verified against a real accessibility audit, not assumed? If yes: use `getByRole`. This is the strongest choice because it's stable against CSS refactors and, as a side effect, it directly validates that the element is usable by assistive technology, which a `data-testid` locator never confirms. Run this check with axe-core or the equivalent accessibility linting your team already has, scoped to the specific screen you're automating, before deciding this is the right locator. Don't assume role-correctness from the framework or component library alone; verify it on the actual rendered page, because custom styling and event handlers frequently break the semantics a library ships by default.
  • If the accessibility tree is incomplete or unreliable for this element, but you can add a test ID: use `getByTestId`. This is the right fallback specifically because it decouples the locator from both the DOM structure and the visible copy, both of which change more often than a deliberately-added test attribute. The tradeoff, and the one most locator guides skip, is the retrofit cost: adding `data-testid` to an existing component means a source change, a PR, a review, and a merge in a codebase you may not fully own. Budget for this as real engineering work, not a QA side task, especially in a legacy app maintained by a different team than the one writing tests.
  • Only when neither is realistically available (a legacy app where markup can't be changed, a third-party embedded widget, an iframe you don't control) does XPath become the appropriate choice, and even then, scope it narrowly: prefer an XPath anchored to stable, semantically meaningful attributes (`//button[@type='submit' and contains(@class, 'checkout')]`) over one anchored to raw DOM position. An XPath that would break if a `<div>` were inserted three levels up is not a fallback, it's a liability with a different syntax.

Tying this to CI stability, not just theory: the practical signal that your locator choice is wrong is not a vague sense of brittleness, it's your own CI flake and failure data. If a specific spec file has a disproportionately high failure/retry rate compared to the rest of the suite, and the failures cluster around one or two locators (visible in Playwright's trace viewer as "element not found" or a strict-mode violation rather than an assertion failure), that's a locator problem, not a flakiness problem, and no amount of retry tuning fixes it. `(Industry consensus)`: teams that track failure reasons per locator type, rather than just pass/fail per test, consistently find that a small number of structurally-anchored selectors account for a disproportionate share of a suite's total flake. Use your own CI history this way before assuming a rewrite is needed: audit which locators are actually failing, not which ones "feel" fragile.

Implementation

getByRole, when the accessibility tree supports it:

import { test, expect } from '@playwright/test';

test('submits the checkout form with valid card details', async ({ page }) => {
 await page.goto('/checkout');

 await page.getByRole('textbox', { name: 'Card number' }).fill('4242424242424242');
 await page.getByRole('textbox', { name: 'Expiry date' }).fill('12/28');
 await page.getByRole('textbox', { name: 'CVC' }).fill('123');

 await page.getByRole('button', { name: 'Place order' }).click();

 await expect(page.getByRole('heading', { name: 'Order confirmed' })).toBeVisible();
});

This only works because the form's inputs and button have correct accessible names, verified in advance, not assumed. If the "Place order" button's visible text changes under an A/B test but its accessible name (via `aria-label`) stays fixed, this locator survives the change. If the accessible name itself is what's under test, add a `data-testid` instead rather than fighting it.

getByTestId, for elements without reliable accessible names:

test('applies a promo code at checkout', async ({ page }) => {
 await page.goto('/checkout');

 await page.getByTestId('promo-code-input').fill('SAVE10');
 await page.getByTestId('promo-code-apply').click();

 await expect(page.getByTestId('order-total')).toContainText('$45.00');
});

Adding these three attributes is the retrofit cost referenced above: three lines of JSX changed, one PR, and a review from the team that owns the checkout component, not a QA-only change.

<input data-testid="promo-code-input" ... />
<button data-testid="promo-code-apply" ...>Apply</button>
<span data-testid="order-total">{formattedTotal}</span>

XPath, scoped narrowly, for legacy markup you cannot change:

test('opens the legacy admin record editor', async ({ page }) => {
 await page.goto('/admin/records/482');

  // Legacy jQuery admin panel, no roles, no test IDs, markup owned by a
  // vendor package that predates the test suite. Anchored to a stable
  // data attribute the vendor markup happens to already emit, not to
  // structural position, so a layout change elsewhere on the page
  // doesn't break this.
 const editButton = page.locator(
    "xpath=//button[@data-action='edit-record' and not(@disabled)]"
  );
 await editButton.click();

 await expect(page.locator("xpath=//div[@id='record-editor-panel']")).toBeVisible();
});

Note what makes this XPath usable rather than a liability: it anchors to a semantic attribute the legacy markup already emits (`data-action`), not to a DOM path like `/html/body/div[4]/div[2]/table/tr[3]`. If the vendor markup emits nothing stable at all, that's the signal to raise the retrofit conversation with whoever owns that surface, rather than writing an increasingly specific structural XPath to compensate.

AI Considerations

An AI coding agent generating Playwright locators from a live page will default to whatever is easiest to extract from the DOM at generation time, which is frequently a CSS selector or a structural path, not the semantically correct locator for that specific element. This matters more than it sounds, because an AI-generated locator that "works" on first run gives no signal about whether it's the resilient choice or the brittle one; both pass the same green CI run today. The practical mitigation is to have the agent (or the human reviewing its output) apply the same three-tier decision check above before accepting a generated locator: does this element have a real role and name, is a test ID already present, and only then is a structural fallback acceptable. An agent instructed to ground every locator against the live accessibility tree, rather than against whatever selector the DOM inspector suggests first, produces meaningfully more durable output, `(Opinion)`, though this depends heavily on the specific tooling and prompting used and is worth validating against your own agent's actual output rather than assumed.

OpenEvident

If you want AI-agent-driven Playwright automation that grounds locators against the live page rather than guessing from a static DOM snapshot, CrevoAI, the Playwright toolkit published under OpenEvident, is built specifically for this workflow: a VS Code or Cursor extension talks to a local MCP server (`runtime-mcp`) that provides codegen and browser-control tools to an AI coding agent (Cursor, Claude Code, GitHub Copilot), while `runtime-worker` drives the actual Playwright/Chromium session, entirely on `127.0.0.1` with no cloud job runner. That local-first architecture matters for the locator-grounding problem specifically, because the agent's browser-control tools operate against the real rendered page rather than a cached or generated approximation of it. `[VERIFY WITH OPENEVIDENT TEAM]`: whether CrevoAI's MCP tools apply any built-in preference ordering between role-based, test-ID, and structural locators, versus this being a decision layer a team still needs to apply itself when reviewing agent-generated tests.

OpenCrevo Implementation

Retrofitting a locator strategy across an existing application, auditing which screens have real accessibility semantics, deciding where `data-testid` is worth the engineering cost, and rewriting the brittle XPath expressions that accumulated along the way, is exactly the kind of test automation maintenance work that's hard to prioritize internally because it competes directly with feature delivery. OpenCrevo's Test Automation service builds repeatable, CI-integrated test suites designed to catch regressions before they reach production, which includes this kind of locator-resilience work as part of the underlying suite quality, not as a separate cleanup project. If this is the piece your team keeps deferring, OpenCrevo can help build it into your existing pipeline. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

This locator work doesn't stand alone. It's the upstream fix for a pattern covered in more depth in Why Your Automated Tests Keep Breaking After Every UI Change, which walks through the before/after of a suite that adopted this same resilience thinking. If your suite's pain point is CI-specific rather than locator-specific (tests that pass locally but fail intermittently in the pipeline), see How to Reduce Flaky Playwright Tests in CI/CD for the retry, isolation, and trace-debugging side of that problem. And if the app you're working with is a legacy Selenium suite rather than a Playwright one, How to Modernize a Legacy Selenium Test Suite Without Starting Again covers the incremental migration path, including how locator strategy fits into that larger move.

Practical Checklist

  • Audit accessibility semantics (axe-core or equivalent) per screen before choosing a locator strategy for it; don't assume role-correctness from the framework alone.
  • Use `getByRole` where the audit confirms a correct, stable role and accessible name.
  • Use `getByTestId` where the accessibility tree is incomplete but a test attribute can realistically be added, and budget the retrofit as real engineering work (PR, review, ownership), not a QA-only task.
  • Reserve XPath for markup you genuinely cannot change, and anchor it to stable semantic attributes, never to raw DOM structure or position.
  • Never use `.first()`, `.last()`, `.nth()`, or `.or()` to paper over a locator that matches more than one element; fix the locator's uniqueness instead.
  • Pull your own CI trace and retry data before rewriting locators speculatively: fix the specific locators your failure data shows are actually breaking, not the ones that feel fragile.
  • Re-run this audit when a design system or component library upgrade changes markup structure or ARIA behavior underneath your existing tests.

FAQ

  • Is XPath always worse than getByRole or data-testid? Not inherently; a narrowly-scoped XPath anchored to a stable attribute is fine. What's actually worse is a structural XPath that depends on DOM position, which is common because it's the easiest kind to write, not because it's the right kind.
  • How long does retrofitting data-testid attributes into a legacy app take? There's no universal number; it scales with how many teams own the components under test and how much review friction exists, `(Opinion)`. Treat it as a per-component engineering task with its own PRs, not a bulk QA change, and prioritize the screens with the highest test-failure rate first.
  • Does getByRole replace the need for a real accessibility audit? No, and this is a common misreading of the advice: `getByRole` locators are only as stable as the accessibility semantics actually are. Using them without first confirming those semantics are correct just relocates the brittleness rather than removing it.
  • What's the difference between a flaky test and a broken locator? A flaky test fails intermittently for timing or environment reasons and often passes on retry. A broken locator fails consistently, or fails specifically when unrelated markup changes; check the failure reason in the trace viewer (not found / strict-mode violation versus a timeout) to tell them apart before treating either as generic flakiness.
  • What breaks first when a team tries to standardize locator strategy across a large legacy app? Usually the assumption that one strategy fits the whole app; the parts with genuine accessibility investment adapt easily to `getByRole`, while older or vendor-owned surfaces need the test-ID or XPath fallback, and treating the whole app as one migration project rather than a per-surface one is where these efforts commonly stall, `(Opinion)`.

Conclusion

The getByRole-versus-data-testid debate isn't wrong, it's just answered one level too early. The real question isn't "which locator type is best," it's "what does this specific screen's accessibility posture actually support, and what does the alternative cost to retrofit." Answer that per component, back it with your own CI failure data instead of a generic brittleness intuition, and reserve XPath for the narrow legacy case it's actually suited to, and the locator strategy stops being a rule copied from a blog post and starts being a decision grounded in the application in front of you.

Sources

Playwright locator documentation, axe-core accessibility testing engine, OpenEvident's `vindicate` repository, industry discussion of locator resilience and coverage-versus-mutation-style tradeoffs in CI stability data (`Industry consensus`, synthesized from current QA-engineering practitioner sources, not a single primary study), OpenCrevo services documentation (`src/data/services.ts` in this repository).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.