Browse by section

QA 日本語

What to Protect with E2E Tests: A Practical Quality Strategy for Playwright and Cypress

The conclusion first: the only things worth protecting with E2E are the boundaries where the result changes your ship-or-hold decision. Not the entire user experience.

Anyone writing E2E tests hits the same wall: how far should E2E reach, what belongs in scope, and what gets cut.

With tools as capable as Playwright and Cypress in wide use, “what is technically writable” and “what should be protected” get conflated easily.

This article covers how to decide what to protect, and how to build quality that works in practice given Playwright or Cypress.

Written for engineers who:

  • Already run E2E tests in CI
  • Have tests, but release decisions have not got easier
  • Agonize over “add or remove” every time

Sponsored

The premise: E2E detects destroyed user value

What E2E protects is that the continuous experience by which a user receives value has not broken.

This is not about tools or schools of thought. Sorting responsibilities by test layer always converges here.

Question Layer that answers it
Is this function correct? Unit tests
Do these modules integrate correctly? Integration tests
Can the user accomplish their goal? E2E tests

So E2E exists to answer “can a user who performs this action reach their goal?”

Drop that premise and E2E becomes a large pile of untrustworthy tests. The reason the test pyramid narrows toward the top is not only that E2E is slow and fragile, but that its granularity is too coarse to decide anything precise.

“It seems important” is a failing criterion

Most teams decide scope like this:

  • It affects revenue, so it matters
  • It is a main screen, so it matters
  • The PM is worried about it, so it matters

I have repeatedly watched suites balloon on that basis until they were useless for deciding anything.

The reason is simple: “seems important” is not usable at the moment something breaks.

What E2E actually needs is not importance but this question:

If this flow breaks, does the ship-or-hold decision change?

Anything you cannot answer instantly is not worth protecting with E2E.

Sponsored

Protect only the boundaries where the decision changes

Protect with E2E Do not protect with E2E
Can users log in? Fine-grained display branching
Can they reach purchase, application, submission? Differences in validation wording
Can the main roles complete critical operations? UI branches that occur under narrow conditions

The right column is not unimportant. It simply is not what E2E decides. Wording belongs in unit tests or snapshots; UI branching belongs in component tests — where a failure points straight at the cause.

The foundation for this is in a quality standard for E2E tests.

How Playwright and Cypress differ as premises

Not “which is better”, but “the constraints differ, so the design does”. Once scope is settled, the implementation is shaped by the tool’s assumptions.

Aspect Playwright Cypress
Where it runs Outside the browser, driving it Inside the browser
Browsers Chromium / Firefox / WebKit Chromium-based / Firefox
Multiple tabs and origins Handled directly Requires cy.origin()
Parallel execution Set a worker count and it runs Official parallelization assumes --record and Cypress Cloud
Debugging Trace Viewer, replayed after the run Test Runner time-travel, during the run

Cross-origin support and parallelization are what actually shape design.

To protect a flow that passes through an external payment screen or SSO, Playwright lets you write it as a single test. Cypress needs cy.origin() to cross origins, which changes how the test is written. Deciding to protect “all the way through payment” turns that difference into implementation cost.

Parallelization is the same. Cypress’s official parallelization is built around recording to Cypress Cloud (--record). Distributing across your own CI alone means splitting specs manually or similar. Playwright takes a worker count, so you can assume “more tests can be absorbed by more parallelism”.

Sponsored

Design principles for either tool

Put “what has to break” in the test name

An E2E test should be explainable in plain language before it is code.

  • Which user
  • Trying to do what
  • Failing where, such that this test fails

If the test name does not convey that, the test is useless for deciding anything.

// Bad: nothing about what breaking causes a failure
test('purchase flow', async ({ page }) => { /* ... */ });

// Good: the moment it fails, you know where to look
test('a logged-in user can reach payment completion from the cart', async ({ page }) => { /* ... */ });

One test, one decision

In either tool, a test that makes several decisions leaves you uncertain when it fails.

Rather than packing login, navigation, input and submission into one test, narrow it to a single “how far does this have to get”.

A sharp test beats a long one. If login is a precondition rather than the thing under test, move it out of the steps with Playwright’s storageState or Cypress’s cy.session() — the intent of the test becomes obvious.

Prefer explainability over stability

E2E does not need to never fail. If a failure immediately tells you two things, occasional failures do not damage quality:

  • What broke
  • Where to start looking

Stability with no explanation is actively harmful. Leave a test that turns green on retry and the team learns that failures can be re-run away. At that point it has stopped being decision material.

Keep Playwright traces, or Cypress screenshots and videos, as CI artifacts, always. Explainability can be bought with configuration.

Naming what you will not protect makes E2E stronger

The most important part of E2E design is deciding what you deliberately will not cover.

Target Where it is protected
Calculation logic, conditional branches Unit tests
API and database integration Integration and API tests
Visual regressions Visual regression testing
Wording and formatting Human review
Whether the goal can be reached E2E

Grow an E2E suite without that split and it collapses, in Playwright or Cypress alike.

The question I always ask:

Looking at this failure, can I decide without hesitating?

Hesitation means the test is badly designed.

For a worked example of how far to take external services like file uploads, see the S3 upload patterns.

Summary: E2E quality is not measured by how much you cover

  • Protect only the boundaries where the decision changes, not the whole experience
  • “Seems important” is not a criterion. “Does a break change the release decision?” is
  • Wording, UI branches and appearance belong to other layers
  • Crossing origins is straightforward in Playwright; Cypress needs cy.origin()
  • Cypress’s official parallelization assumes Cypress Cloud; Playwright just takes a worker count
  • Put “what has to break” into the test name
  • One test, one decision. Move login into storageState or cy.session()
  • Explainability over stability. Keep traces and videos as CI artifacts
  • Tolerating tests that go green on retry is what really lowers quality

Playwright and Cypress are instruments for checking those boundaries quickly and reliably. The tool does not decide the scope.

You will know the quality of an E2E suite not when the test count rises, but when release decisions get faster.