The conclusion first: the only things worth protecting with E2E are the boundaries where the result changes your ship-or-hold decision. Not the entire user experience.
Anyone writing E2E tests hits the same wall: how far should E2E reach, what belongs in scope, and what gets cut.
With tools as capable as Playwright and Cypress in wide use, “what is technically writable” and “what should be protected” get conflated easily.
This article covers how to decide what to protect, and how to build quality that works in practice given Playwright or Cypress.
Written for engineers who:
- Already run E2E tests in CI
- Have tests, but release decisions have not got easier
- Agonize over “add or remove” every time
Sponsored
The premise: E2E detects destroyed user value
What E2E protects is that the continuous experience by which a user receives value has not broken.
This is not about tools or schools of thought. Sorting responsibilities by test layer always converges here.
| Question | Layer that answers it |
|---|---|
| Is this function correct? | Unit tests |
| Do these modules integrate correctly? | Integration tests |
| Can the user accomplish their goal? | E2E tests |
So E2E exists to answer “can a user who performs this action reach their goal?”
Drop that premise and E2E becomes a large pile of untrustworthy tests. The reason the test pyramid narrows toward the top is not only that E2E is slow and fragile, but that its granularity is too coarse to decide anything precise.
“It seems important” is a failing criterion
Most teams decide scope like this:
- It affects revenue, so it matters
- It is a main screen, so it matters
- The PM is worried about it, so it matters
I have repeatedly watched suites balloon on that basis until they were useless for deciding anything.
The reason is simple: “seems important” is not usable at the moment something breaks.
What E2E actually needs is not importance but this question:
If this flow breaks, does the ship-or-hold decision change?
Anything you cannot answer instantly is not worth protecting with E2E.
Sponsored
Protect only the boundaries where the decision changes
| Protect with E2E | Do not protect with E2E |
|---|---|
| Can users log in? | Fine-grained display branching |
| Can they reach purchase, application, submission? | Differences in validation wording |
| Can the main roles complete critical operations? | UI branches that occur under narrow conditions |
The right column is not unimportant. It simply is not what E2E decides. Wording belongs in unit tests or snapshots; UI branching belongs in component tests — where a failure points straight at the cause.
The foundation for this is in a quality standard for E2E tests.
How Playwright and Cypress differ as premises
Not “which is better”, but “the constraints differ, so the design does”. Once scope is settled, the implementation is shaped by the tool’s assumptions.
| Aspect | Playwright | Cypress |
|---|---|---|
| Where it runs | Outside the browser, driving it | Inside the browser |
| Browsers | Chromium / Firefox / WebKit | Chromium-based / Firefox |
| Multiple tabs and origins | Handled directly | Requires cy.origin() |
| Parallel execution | Set a worker count and it runs | Official parallelization assumes --record and Cypress Cloud |
| Debugging | Trace Viewer, replayed after the run | Test Runner time-travel, during the run |
Cross-origin support and parallelization are what actually shape design.
To protect a flow that passes through an external payment screen or SSO, Playwright lets you write it as a single test. Cypress needs cy.origin() to cross origins, which changes how the test is written. Deciding to protect “all the way through payment” turns that difference into implementation cost.
Parallelization is the same. Cypress’s official parallelization is built around recording to Cypress Cloud (--record). Distributing across your own CI alone means splitting specs manually or similar. Playwright takes a worker count, so you can assume “more tests can be absorbed by more parallelism”.
Sponsored
Design principles for either tool
Put “what has to break” in the test name
An E2E test should be explainable in plain language before it is code.
- Which user
- Trying to do what
- Failing where, such that this test fails
If the test name does not convey that, the test is useless for deciding anything.
// Bad: nothing about what breaking causes a failure
test('purchase flow', async ({ page }) => { /* ... */ });
// Good: the moment it fails, you know where to look
test('a logged-in user can reach payment completion from the cart', async ({ page }) => { /* ... */ });
One test, one decision
In either tool, a test that makes several decisions leaves you uncertain when it fails.
Rather than packing login, navigation, input and submission into one test, narrow it to a single “how far does this have to get”.
A sharp test beats a long one. If login is a precondition rather than the thing under test, move it out of the steps with Playwright’s storageState or Cypress’s cy.session() — the intent of the test becomes obvious.
Prefer explainability over stability
E2E does not need to never fail. If a failure immediately tells you two things, occasional failures do not damage quality:
- What broke
- Where to start looking
Stability with no explanation is actively harmful. Leave a test that turns green on retry and the team learns that failures can be re-run away. At that point it has stopped being decision material.
Keep Playwright traces, or Cypress screenshots and videos, as CI artifacts, always. Explainability can be bought with configuration.
Naming what you will not protect makes E2E stronger
The most important part of E2E design is deciding what you deliberately will not cover.
| Target | Where it is protected |
|---|---|
| Calculation logic, conditional branches | Unit tests |
| API and database integration | Integration and API tests |
| Visual regressions | Visual regression testing |
| Wording and formatting | Human review |
| Whether the goal can be reached | E2E |
Grow an E2E suite without that split and it collapses, in Playwright or Cypress alike.
The question I always ask:
Looking at this failure, can I decide without hesitating?
Hesitation means the test is badly designed.
For a worked example of how far to take external services like file uploads, see the S3 upload patterns.
Summary: E2E quality is not measured by how much you cover
- Protect only the boundaries where the decision changes, not the whole experience
- “Seems important” is not a criterion. “Does a break change the release decision?” is
- Wording, UI branches and appearance belong to other layers
- Crossing origins is straightforward in Playwright; Cypress needs
cy.origin() - Cypress’s official parallelization assumes Cypress Cloud; Playwright just takes a worker count
- Put “what has to break” into the test name
- One test, one decision. Move login into
storageStateorcy.session() - Explainability over stability. Keep traces and videos as CI artifacts
- Tolerating tests that go green on retry is what really lowers quality
Playwright and Cypress are instruments for checking those boundaries quickly and reliably. The tool does not decide the scope.
You will know the quality of an E2E suite not when the test count rises, but when release decisions get faster.
