Browse by section

QA 日本語

What Makes a High-Quality E2E Test? A Practical Quality Standard Based on Global Models and Real-World Experience

The conclusion first: the quality standard for an E2E test is whether that test pushes a decision in the right direction. Not coverage, and not a green CI run.

Search for E2E test quality standards and you will always find the same list: stable, covers the critical flows, automated in CI. None of that is wrong.

But I have repeatedly seen E2E suites that satisfy every one of those and are still worthless in practice. The tests pass. CI is green. Production incidents happen anyway, and people still agonize over whether a change is safe to ship.

This article first sets out the internationally shared view of software quality, then covers the gap that opens when you apply it literally, and what a genuinely useful standard looks like.

Sponsored

What the “international standard” says about quality

There is a globally referenced framework for software quality: ISO/IEC 25010.

It was revised in 2023, and now defines nine quality characteristics — up from eight in the 2011 edition. Older articles still describe the old set, so check the date on anything you read.

ISO/IEC 25010:2023 characteristic Meaning
Functional suitability Does it behave as expected?
Performance efficiency Is performance appropriate for the resources used?
Compatibility Can it coexist and interoperate with other systems?
Interaction capability Is it understandable and usable? Replaces “usability” from 2011
Reliability Does it keep working stably?
Security Does it protect information and access?
Maintainability Is it easy to change and fix?
Flexibility Can it adapt to changing environments? Replaces “portability”
Safety Does it avoid harm to people and assets? New in 2023

The important part is that “test quality” is itself evaluated as a means of protecting these characteristics. A test has no value on its own terms.

Under that framing, an E2E suite is:

A mechanism for continuously judging whether the whole system is correct, stable, and safe to use from a user’s point of view

So far, entirely reasonable. The problem starts when you carry that straight onto a real team.

Applying the standard literally breaks your E2E suite

The ISO model is correct. It is also far too abstract.

Translated onto a team, it usually becomes:

  • Every critical user flow is protected by E2E
  • Automate as many cases as possible
  • Keep CI green and stable at all times

None of that looks wrong. And yet I have watched suites written on exactly that basis become useless in practice, more than once.

The reason is simple: “how we measure quality” and “whether this helps anyone decide anything” have come apart.

Take “every critical flow is protected by E2E” literally and the suite grows without bound, run times stretch, flaky tests creep in, and eventually you reach the state where nobody is surprised when it fails. Reliability in the model’s terms may still be satisfied; as decision material, it is dead.

Sponsored

The core: E2E tests do not measure quality

This is where things go wrong most often.

E2E tests are not for quantifying or exhaustively covering quality. They are material for making a decision about quality.

  • Is this change safe to release?
  • What broke?
  • How far should I suspect the blast radius extends?

Push the quality model far enough and the question that remains is:

Looking at this test result, can a person make the right call?

A suite that fails this lowers quality in practice, however well it appears to satisfy the model on paper. It delays release decisions and eats investigation time.

Why a seemingly subjective standard is necessary

Subjectivity enters here, but this is not sentiment.

E2E failures usually show up like this:

What happens Consequence
The test failed, but the cause is not obvious Investigation eats the day
Nobody can tell whether something is actually broken The team learns to re-run and ignore a green retry
Someone ends up reading the code anyway Running the suite stops meaning anything

At that moment, the suite stops helping decisions and starts delaying them.

So the standard I use is:

When it breaks, is there any hesitation?

That is not a feeling — it is the quality of the decision, which is about as fundamental as this gets. In ISO/IEC 25010 terms, it asks about the maintainability of the test code and the interaction capability of its output: how legible the result is to the person reading it.

Sponsored

A quality standard you can actually use

Grounded in both the model and experience, this is the standard I still use.

Can you answer these three immediately?

  1. What has to break for this test to fail?
  2. When it fails, is the next action obvious?
  3. Does this result strengthen the release decision?

These are the ISO characteristics pushed all the way down to the working level.

If even one of them cannot be answered on the spot, that test is probably dulling decisions while appearing to protect quality.

How to apply them

Situation Use
Writing a new test Answer 1-3 before writing. If you cannot, do not write it
Auditing an existing suite Tests that fail question 1 are the first deletion candidates
Handling a flaky test If question 2 has no answer, fix it or delete it. Leaving it is the worst option
Review Ask “what has to break for this to fail?” No answer means the design is vague

Failing the first question is the most dangerous case. A single test that logs in, adds to cart, checks out and more gives you too many suspects to narrow down when it fails.

How to choose what belongs in E2E at all is covered in what to protect with E2E tests.

Summary: the real standard is decision accuracy

  • The quality standard for an E2E test is whether it strengthens a correct decision
  • ISO/IEC 25010 has nine characteristics since the 2023 revision (safety added; usability became interaction capability; portability became flexibility)
  • The model is right but too abstract; applied literally, the suite bloats
  • E2E tests are decision material, not a measuring instrument
  • A suite you cannot reason about shifts from helping decisions to delaying them
  • Three questions: what breaks it, what you do when it breaks, whether it strengthens release decisions
  • The sprawling test that fails question one is the worst offender

None of this departs from the international standard. It is the abstract model brought back into a usable shape.

  • Meaning over coverage
  • Explanatory power over stability
  • Decisions over counts

The quality of an E2E suite does not live in the test code. It shows up in people’s decisions.