Skip to main content
Test AutomationUI ChangesE2E Testing

Why Your Automated Tests Keep Breaking After Every UI Change, and How to Fix It

2 September 2026 · OpenCrevo

A designer renames a button from "Submit" to "Continue," or a developer wraps a form field in a new `<div>` for styling reasons, and suddenly a dozen automated tests turn red. Nothing about the actual feature broke. The tests broke. If this pattern feels familiar, you're not dealing with bad luck or an unusually unstable product, you're dealing with a locator strategy that was never designed to survive change in the first place. This article explains why tests break this way, walks through a real before/after example of a brittle test versus a resilient one, and gives you a practical way to tell, before you write the next test, whether it will survive the next redesign.

The Real Problem

Here's the pattern almost every team runs into within a few months of adopting UI test automation: the suite passes reliably for a while, then a routine front-end change, a rename, a restyle, a component library upgrade, a new wrapper element, causes several tests to fail at once. The team investigates, finds nothing wrong with the product, updates the selectors, and moves on. Then it happens again next sprint. Eventually someone asks the question that actually matters: why does a cosmetic change keep breaking tests that were supposed to verify behavior, not markup?

The honest answer is usually that the tests were never verifying behavior in the first place. They were verifying markup, and markup is the least stable part of a modern web application. A locator like `div.container > div:nth-child(3) > button.btn-primary` is not describing "the submit button," it's describing a specific structural path through the DOM as it happened to exist on the day the test was written. Change the structure even slightly (a wrapping `<div>` for a new layout, a class rename from a CSS refactor, a reordered element) and the path is gone, even though the button itself is still exactly where a real user would expect it.

This is the core insight this article is built around: tests don't break because UI changes are frequent, they break because the locator choices made when the test was written were tied to implementation details instead of to what a user actually sees and interacts with. Fixing the symptom (updating selectors after each break) never stops the pattern. Fixing the cause means changing how locators are chosen before the test is ever written.

Why This Happens

A few things combine to make this the default outcome rather than an edge case:

  • CSS selectors and browser DevTools make the brittle path the easy path. Right-click an element, "Copy selector," and you get something like `#root > div > div.MuiBox-root.css-1a2b3c > button`. It works immediately, so it feels productive. It also encodes a CSS framework's auto-generated class name and the exact nesting depth at that moment, both of which are implementation details with no contract to stay stable.
  • Structural coupling, not visual coupling. A CSS-path or deep-nesting selector is coupled to how the DOM is built, not to how the page looks or behaves. A visual redesign that keeps the same user-facing behavior (same button, same label, same position) can still restructure the underlying markup completely, because there's no reason a front-end developer would think to preserve DOM structure for the benefit of a test suite they may not even know exists.
  • Class names churn faster than anything else in modern front-end code. CSS-in-JS libraries, utility-class frameworks, and component library upgrades routinely regenerate class names on every build. A selector anchored to a class name is effectively anchored to a build artifact, not a stable identifier `(Industry consensus)`.
  • Nobody budgets time for locator maintenance until it's already a fire. Teams plan for writing new tests, rarely for the fact that a chosen locator strategy has an ongoing cost every time the UI changes. That cost is invisible until the first UI refresh, at which point it looks like "the tests are just flaky" rather than "we picked an unstable identifier."

Common Approaches That Fail

  • Fixing selectors reactively, one break at a time. This treats the symptom every single time and never reduces the rate of future breakage, because the next UI change hits the same category of fragile selector somewhere else in the suite.
  • Freezing the UI to protect the tests. Occasionally proposed, never actually adopted, and backwards: the tests exist to serve the product, not the other way around.
  • Adding `.first()`, `.nth()`, or broader wildcard selectors to "make it pass again." This often just trades one fragile match for a worse one: a selector that now matches multiple elements and silently picks the wrong one after the next change, producing a test that passes for the wrong reason instead of failing honestly.
  • Blaming the testing framework. Switching from one framework to another (Selenium to Cypress, Cypress to Playwright) does not fix a fragile-locator problem, because the fragility lives in the selector strategy, not the tool driving the browser.
  • "Just add more tests to catch what we missed." More tests written with the same brittle locator approach simply means more tests break on the next UI change, at higher maintenance cost, not better coverage.

Practical Solution

The fix is choosing locators anchored to what a user actually perceives and interacts with (a role, a label, a stable identifier deliberately placed in the markup for this purpose), instead of locators anchored to the DOM's incidental structure. Concretely, that means a locator priority order, roughly in this sequence:

  • Accessible role and name (`getByRole('button', { name: 'Continue' })`), because this reflects what a screen reader and a real user both perceive, and a11y roles are far less likely to change on a purely cosmetic refactor.
  • A dedicated test identifier (`data-testid`, or an equivalent attribute), deliberately added to markup specifically so tests have something stable to hold onto, independent of styling or layout.
  • Visible text, when the role alone isn't unique enough, understanding that copy changes are still a risk, just a smaller and more visible one than a structural change.
  • CSS structural selectors and `nth-child` chains, only as a last resort, and treated as technical debt the moment they're written, not a first choice.

The before/after example below is the concrete version of the content gap this article exists to close: it isn't enough to say "tests are brittle," the connection between locator choice and survivability needs to be shown directly.

Before: a brittle, CSS-path-based test

// checkout.spec.ts (brittle version)
import { test, expect } from '@playwright/test';

test('user can submit checkout form', async ({ page }) => {
 await page.goto('/checkout');

  // Selector copied straight from DevTools "Copy selector"
 await page.fill('#root > div > div.MuiBox-root.css-1a2b3c > form > div:nth-child(2) > input', 'jane@example.com');

 await page.click('#root > div > div.MuiBox-root.css-1a2b3c > form > div.btn-wrapper > button.btn-primary');

 await expect(page.locator('div.confirmation-banner')).toBeVisible();
});

Now suppose the design team wraps the checkout form in a new responsive layout container for a mobile redesign, and the CSS-in-JS build regenerates `css-1a2b3c` to a new hash on the next deploy, as CSS-in-JS libraries routinely do. None of the actual behavior changed. The email field still exists, the submit button still says the same thing, the confirmation banner still appears. But every selector in this test references a DOM path and a class name that no longer exist, so the test fails on `page.fill(...)` before it even reaches the assertion that matters.

After: a resilient, role-and-testid-based test

// checkout.spec.ts (resilient version)
import { test, expect } from '@playwright/test';

test('user can submit checkout form', async ({ page }) => {
 await page.goto('/checkout');

 await page.getByRole('textbox', { name: 'Email' }).fill('jane@example.com');

 await page.getByRole('button', { name: 'Continue' }).click();

 await expect(page.getByTestId('checkout-confirmation')).toBeVisible();
});

This version survives the same redesign scenario. The wrapping container, the regenerated CSS-in-JS class name, and the reordered `div`s are all invisible to this test, because none of the locators reference DOM structure or class names at all. `getByRole` matches on the accessible role and visible label, which a purely cosmetic markup change does not touch. `getByTestId` matches a `data-testid="checkout-confirmation"` attribute that a developer deliberately added to that element specifically so it stays addressable regardless of styling changes: `<div data-testid="checkout-confirmation" className={styles.banner}>`.

The difference isn't "this test is written more carefully." It's that the before version encodes an assumption ("the DOM will look exactly like this forever") that nothing in a normal front-end workflow promises to keep true, and the after version encodes an assumption ("this button will still be labeled Continue and reachable as a button") that a routine restyle has no reason to break.

Implementation

Rolling this out across an existing suite doesn't require rewriting everything at once. A practical sequence:

  • Audit the current suite for CSS-path and `nth-child` selectors. A quick grep across the test directory for patterns like `nth-child`, auto-generated class prefixes (`css-`, `MuiBox`, `sc-` for styled-components), and multi-level `>` combinators will surface most of the risk quickly.
  • Prioritize by breakage frequency, not by file order. The tests that have broken the most times in the last few months are the ones with the worst locators; fix those first for the fastest return.
  • Add `data-testid` attributes to elements that don't have a strong accessible role, in collaboration with front-end developers, since this is a one-line markup addition (`data-testid="checkout-confirmation"`) that costs almost nothing to add but removes an entire category of future breakage.
  • Rewrite failing tests with role-first locators as they break, rather than doing a big-bang rewrite of a suite that's currently passing. There's no value in touching tests that aren't failing yet.
  • Add a lint rule or code-review checklist item that flags new CSS-path selectors in pull requests, so the fix doesn't erode back to old habits within a few sprints. `[VERIFY API BEFORE PUBLICATION]`: specific ESLint plugin rule names for Playwright locator linting should be confirmed against current plugin versions before including a exact rule name in a published draft.

For teams evaluating locator strategy in more depth, including the getByRole-versus-`data-testid` tradeoff and a full decision framework, see our companion article, Stop Using Fragile CSS Selectors: A Practical Locator Strategy.

AI Considerations

AI coding agents that generate Playwright tests inherit this exact problem, often at a larger scale than a human would produce manually, because an agent asked to "write a test for this page" without explicit instruction will frequently reach for whatever selector is easiest to derive from the current DOM snapshot, which is often a CSS path (Industry consensus). If you're using an AI agent to scaffold or extend a test suite, the locator-priority order above (role, then test-id, then text, then CSS as a last resort) is worth stating explicitly in the agent's instructions or project rules, rather than assuming the agent will infer it. An agent that is told to ground every locator against the live page, and to prefer role-based or test-id-based locators, produces materially more durable tests than one left to guess from a static markup read.

This also connects to a second, related failure mode worth flagging separately: a suite that breaks constantly on every UI change is also a suite where flaky-looking failures (tests that fail intermittently for reasons unrelated to a real UI change, such as timing or environment differences) get harder to distinguish from genuine locator breakage, because both show up as the same red build. Our companion article, How to Reduce Flaky Playwright Tests in CI/CD, covers that adjacent problem in depth.

OpenEvident

If you're already using an AI coding agent (Cursor, Claude Code, GitHub Copilot) to write or maintain Playwright tests, the locator-grounding discipline described above is easier to enforce consistently with tooling built for that workflow specifically. CrevoAI, the Playwright test automation toolkit published under OpenEvident, is described in its own documentation as "AI-native Playwright test automation" with MCP tools for codegen and browser control, and it runs as a fully local stack (a VS Code/Cursor extension talking to a local MCP server and worker, with "no cloud job machine, no MongoDB, no remote orchestration" per its architecture docs). That local, agent-facing architecture matters here specifically because it gives an AI agent a live, grounded view of the actual page during codegen, rather than generating locators from a stale or incomplete snapshot. `[VERIFY WITH OPENEVIDENT TEAM]`: whether CrevoAI's codegen tooling enforces or defaults to a specific locator priority order (role-first, test-id-second) versus leaving that choice to the agent's own instructions.

OpenCrevo Implementation

Auditing and re-anchoring locators across an existing, large test suite is exactly the kind of project that's simple in a single test file and genuinely time-consuming across hundreds of them, especially when it also means coordinating with front-end teams on adding `data-testid` attributes. OpenCrevo's Test Automation service focuses on repeatable, CI-integrated test suites built to catch regressions before they reach production, which includes this kind of locator-resilience work as part of building a suite meant to last. If untangling a brittle suite across an entire application is the part that's stalled your team, OpenCrevo can help build that transformation without asking you to freeze feature work while it happens. Not sure where your gaps are? Start with the free QA maturity assessment for a scored baseline before scoping the engagement.

Practical Checklist

  • Grep your test suite for `nth-child`, CSS-in-JS class prefixes (`css-`, `MuiBox`, `sc-`), and multi-level `>` selector chains, and treat every match as a maintenance liability.
  • Default new locators to `getByRole` with an accessible name first, `data-testid` second, visible text third, CSS structural selectors only as a documented last resort.
  • Add `data-testid` attributes to elements during development, not retroactively after the third test failure references them.
  • Fix the highest-breakage-frequency tests first; don't rewrite a suite that isn't currently failing.
  • Add a lint rule or PR checklist item that catches new CSS-path selectors before merge, so the fix holds after the initial cleanup.
  • If using an AI agent to generate tests, state the locator priority order explicitly in its instructions rather than assuming it will infer one.

FAQ

  • Is switching locator strategy worth it for a small team with a small suite? Yes, arguably more so: a small team has less capacity to absorb repeated selector-fixing sprints, and the fix (role-first, test-id-second) costs the same whether you have 10 tests or 1,000.
  • How long does re-anchoring an existing suite take? For a single file, minutes. For a full legacy suite, it depends entirely on size, but prioritizing by breakage frequency (fixing the worst offenders first) delivers most of the benefit long before the whole suite is done.
  • Does this fully replace the need for a human to update tests after real feature changes? No. A locator strategy anchored to role or test-id survives cosmetic and structural changes, not changes to what the feature actually does; if a button is genuinely removed or its behavior changes, the test should still need updating, and that's a correct failure, not a brittle one.
  • What's the difference between this problem and flaky tests? They look similar (a red build) but have different causes: a brittle-locator failure happens deterministically after a specific UI change and stays broken until fixed, while a flaky test fails intermittently for reasons often unrelated to any real change, such as timing. Both matter, but they need different diagnoses.
  • Does adding `data-testid` attributes clutter production markup? It's a small, inert HTML attribute with no runtime cost and no effect on styling or behavior; the tradeoff is a negligible amount of extra markup in exchange for meaningfully fewer broken builds.

Conclusion

A test suite that goes red every time someone touches the UI isn't telling you the product is unstable, it's telling you the locators were never anchored to anything the UI actually promises to keep stable. The fix isn't a bigger regression suite or a different test runner, it's choosing locators (role first, test-id second, text third, CSS structure only as a last resort) that describe what a user sees and does, not the DOM path that happened to exist on the day the test was written. Make that choice upfront, on the next test you write, and the next redesign becomes a non-event instead of a fire drill.

Sources

Playwright locator documentation on getByRole and getByTestId, industry discussion of CSS-in-JS class-name instability and structural selector fragility (Industry consensus, synthesized from multiple current test-automation engineering publications, not a single primary source), OpenEvident's vindicate repository and architecture documentation, OpenCrevo services documentation (src/data/services.ts in this repository).

START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.