Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

Flaky Tests Hide Behind Retries: How to See Them and Quarantine Them Properly

What a retried flaky test actually looks like in Playwright's list, JUnit, and JSON reporters, how we diagnosed a real one on this site, and a quarantine policy with owners and expiry dates that CI enforces.

QA Vibes EditorialPublished Updated 11 minTested with Playwright 1.63.0, Node.js 22.16.0Revision history ↓

Key takeaways

  • A test that passes on retry leaves no trace in Playwright's JUnit report; add the JSON reporter to see flakes.
  • Measure how often a test flakes before changing anything, and change one factor at a time.
  • Quarantine a test you can't fix today with an owner, an issue, and an expiry date that CI enforces.
Contents (8 sections)

Introduction

A flaky test is one that fails and then passes without any code change. With retries enabled, Playwright runs a failed test again, and if the second attempt passes, the run can still exit green. That is useful: one slow network call shouldn't block a release. It's also dangerous: the pipeline stays green while a real race condition sits in the product or the test.

This guide covers three things, all checked on Playwright 1.63.0:

  1. What a retried flaky test actually looks like in each reporter, including one that records nothing at all.
  2. How we diagnosed a real flaky test on this site, with the numbers.
  3. A quarantine policy with owners and expiry dates, enforced by a short script in CI.

What a retried flake looks like in each reporter

We built a small project with three tests and retries: 1: one stable test, one that fails on its first attempt and passes on the retry, and one quarantined test we'll use later. The flaky test is deterministic, so every run shows the same result:

import { test, expect } from "@playwright/test";
 
// Deterministically flaky: fails on the first attempt, passes on the retry.
test("checkout button navigates", async ({ browserName }, testInfo) => {
  expect(testInfo.retry, "simulated flake on the first attempt").toBeGreaterThan(0);
});

The config writes three reports at once:

import { defineConfig } from "@playwright/test";
 
export default defineConfig({
  testDir: "./tests",
  retries: 1,
  reporter: [
    ["list"],
    ["json", { outputFile: "results/results.json" }],
    ["junit", { outputFile: "results/results.xml" }],
  ],
});

Running the stable and flaky tests gave this:

Where you look What it showed
Exit code 0: the run passed
List reporter (terminal) 1 flaky, 1 passed
JUnit XML tests="2" failures="0", no retry, no flaky marker
JSON report "flaky": 1, test status "flaky", attempts failed then passed
Exit code with --fail-on-flaky-tests 1: the run failed

The JUnit report records a clean pass

This is the JUnit XML from that run, unedited:

<testsuites id="" name="" tests="2" failures="0" skipped="0" errors="0" time="0.675869">
<testsuite name="flaky.spec.ts" timestamp="2026-09-13T19:50:47.230Z" hostname="" tests="1" failures="0" skipped="0" time="0.008" errors="0">
<testcase name="checkout button navigates" classname="flaky.spec.ts" time="0.008">
<system-out>
<![CDATA[
[[ATTACHMENT|..\..\..\test-results\flaky-checkout-button-navigates\error-context.md]]
]]>
</system-out>
</testcase>
</testsuite>
<testsuite name="stable.spec.ts" timestamp="2026-09-13T19:50:47.230Z" hostname="" tests="1" failures="0" skipped="0" time="0.004" errors="0">
<testcase name="cart total adds up" classname="stable.spec.ts" time="0.004">
</testcase>
</testsuite>
</testsuites>

The flaky test has no <failure> element, no retry count, and no flaky marker. The only hint is an attachment path to the failed attempt's error context. A dashboard or test management tool that reads only this file sees two passing tests.

This is a known limitation, not a bug. A 2025 request to add an opt-in flaky marker to the JUnit reporter was closed with a "collecting feedback" label. Maintainers have preferred to keep that reporter standard JUnit.

The JSON report records the truth

Trimmed to the relevant fields, the same run's JSON report looks like this:

{
  "stats": { "expected": 1, "skipped": 0, "unexpected": 0, "flaky": 1 },
  "tests": [
    {
      "title": "checkout button navigates",
      "status": "flaky",
      "attempts": [
        { "retry": 0, "status": "failed" },
        { "retry": 1, "status": "passed" }
      ]
    }
  ]
}

In the real report, tests are nested under suites[].specs[].tests[], and each attempt lives in results[]. The script later in this guide walks that structure.

Step 1: Make flakes visible

Keep retries if you need them, but stop letting them hide flakes:

  • Add the JSON reporter next to whatever reporter feeds your dashboard. The terminal output says a test was flaky, but JSON is the report file a script can read that still says so.
  • Fail the run on flakes with --fail-on-flaky-tests. The retry still tells you whether the failure was persistent, but the pipeline goes red either way.
  • Record traces on the first retry with trace: "on-first-retry" in the use block of your config. Every flake then leaves a trace you can open in the trace viewer.
npx playwright test --fail-on-flaky-tests

Step 2: Reproduce before you fix

This site's own test suite had a flaky test: the smoke test that clicks through every link in the main navigation. Here is how we found the cause, including two mistakes along the way.

Measure a rate, not an anecdote. A single failure proves nothing. We ran the test repeatedly with the same parallelism as a normal run:

npx playwright test tests/e2e/smoke.spec.ts --grep "main navigation" --repeat-each=20 --workers=6 --output=.flake-evidence

It failed 3 times in 15. The --output folder matters: every new run clears the default test-results folder, and our first round of evidence was deleted exactly that way.

Mistake 1: a fix that only moved the failure. The test waited for a heading to be visible before checking which link was active. We assumed a timing race and made it wait for the URL instead. Failures dropped but didn't stop. The traces showed why: the click completed, the page data request returned 200 in 16 to 191 ms, and the page still never changed. Nothing was slow. The navigation simply never committed.

Mistake 2: a measurement that ran zero tests. To isolate the cause, we wrote a small experiment and filtered it with --grep. The filter matched nothing, Playwright reported "No tests found", and our script counted that as zero failures. Always check the number of tests a run executed before trusting its failure count.

Change one factor at a time. The experiment clicked a second link while the first navigation was still finishing:

Configuration Runs Second click lost
Next.js 16.0.3, 6 workers, header links 40 8
Next.js 16.0.3, 6 workers, footer links with no click handler 40 5
Next.js 16.0.3, 1 worker 80 0
Next.js 16.3.5, 6 workers 80 0

Footer links failed too, which cleared our own menu code. With a single worker, nothing failed, so the behavior needed CPU contention to appear. Upgrading the framework removed it under the same load, and the smoke test then passed 40 out of 40 runs on desktop and mobile. No release note names this fix, so we only claim what we measured on our machine. It's also why we now pin exact framework versions.

What we did not do: add a longer wait, add a retry, or mark the test fixme. Each of those would have turned a measurable problem into an invisible one. Our CI retries failed tests once, so there this flake would usually have passed on the retry and shown up as green.

Many flakes have more ordinary causes: fixed waits, missing await, or assertions that don't retry. Our Playwright test reviewer checks a pasted test for those patterns.

Step 3: Quarantine with an owner and an expiry date

Sometimes a flaky test can't be fixed today. Quarantine lets the rest of the suite keep blocking merges while that one test stops blocking. Done badly, quarantine becomes a permanent skip list. Done well, every quarantined test still runs, has an owner, links to an issue, and expires.

Tag and annotate the test

Playwright tags and annotations carry that information in the test itself:

import { test, expect } from "@playwright/test";
 
test(
  "search suggestions appear",
  {
    tag: "@quarantine",
    annotation: {
      type: "quarantine",
      description: "owner=@qa-lead issue=https://github.com/example/shop/issues/482 expires=2026-10-15",
    },
  },
  async () => {
    expect("no suggestions").toBe("3 suggestions");
  },
);

--grep-invert @quarantine runs everything except quarantined tests, and --grep @quarantine runs only them. Both tags and annotations appear in the JSON report. One detail cost us a few minutes: the report stores the tag as quarantine, without the @, so a script that looks for "@quarantine" never matches.

Enforce the policy with the JSON report

This script reads the JSON report. It prints a warning for each flaky test, notes quarantined tests that pass again, and fails if a quarantine is missing its owner, issue, or expiry date, or has expired. On GitHub Actions it also adds a table of flaky tests to the job summary.

// flake-report.mjs: reads Playwright's JSON report, lists flaky tests, and enforces a
// quarantine policy. Every @quarantine test needs an owner, an issue, and an expiry date.
// Usage: node flake-report.mjs results/results.json
import fs from "node:fs";
 
const report = JSON.parse(fs.readFileSync(process.argv[2] ?? "results/results.json", "utf8"));
const today = new Date().toISOString().slice(0, 10);
const flaky = new Map();
const problems = new Set();
const recovered = new Set();
 
function visit(suite) {
  for (const spec of suite.specs ?? []) {
    for (const test of spec.tests) {
      const name = `${spec.file} › ${spec.title}`;
      if (test.status === "flaky") {
        flaky.set(name, test.results.map((r) => r.status).join(" → "));
      }
      // Playwright's JSON report stores tags without the leading "@".
      if (!spec.tags.includes("quarantine")) continue;
 
      const note = test.annotations.find((a) => a.type === "quarantine")?.description ?? "";
      const owner = note.match(/owner=(\S+)/)?.[1];
      const issue = note.match(/issue=(\S+)/)?.[1];
      const expires = note.match(/expires=(\d{4}-\d{2}-\d{2})/)?.[1];
 
      if (!owner || !issue || !expires) {
        problems.add(`${name}: quarantine annotation needs owner=, issue=, and expires=YYYY-MM-DD`);
      } else if (expires < today) {
        problems.add(`${name}: quarantine expired on ${expires} (owner ${owner}, ${issue})`);
      }
      if (test.status === "expected") recovered.add(name);
    }
  }
  for (const child of suite.suites ?? []) visit(child);
}
 
report.suites.forEach(visit);
 
for (const [name, attempts] of flaky) console.log(`::warning title=Flaky test::${name} (${attempts})`);
for (const name of recovered) console.log(`::notice title=Quarantined test passed::${name}. If it keeps passing, take it out of quarantine.`);
for (const problem of problems) console.log(`::error title=Quarantine policy::${problem}`);
 
if (process.env.GITHUB_STEP_SUMMARY) {
  const rows = [...flaky].map(([name, attempts]) => `| ${name} | ${attempts} |`).join("\n");
  const table = flaky.size ? `| Test | Attempts |\n| --- | --- |\n${rows}\n` : "None this run.\n";
  fs.appendFileSync(process.env.GITHUB_STEP_SUMMARY, `### Flaky tests: ${flaky.size}\n\n${table}`);
}
 
console.log(`${flaky.size} flaky, ${recovered.size} recovered in quarantine, ${problems.size} policy problem(s)`);
process.exit(problems.size > 0 ? 1 : 0);

We ran it against the real reports from our example project, plus copies edited to cover each failure case:

flaky run          ::warning title=Flaky test::flaky.spec.ts › checkout button navigates (failed → passed)
                   1 flaky, 0 recovered in quarantine, 0 policy problem(s)             exit 0
valid quarantine   0 flaky, 0 recovered in quarantine, 0 policy problem(s)             exit 0
expired            ::error title=Quarantine policy::quarantine.spec.ts › search suggestions appear: quarantine expired on 2026-09-01 (owner @qa-lead, https://github.com/example/shop/issues/482)
                   exit 1
missing owner      ::error title=Quarantine policy::quarantine.spec.ts › search suggestions appear: quarantine annotation needs owner=, issue=, and expires=YYYY-MM-DD
                   exit 1
passing again      ::notice title=Quarantined test passed::quarantine.spec.ts › search suggestions appear. If it keeps passing, take it out of quarantine.
                   exit 0

On GitHub Actions, the ::warning, ::notice, and ::error lines become annotations on the run.

Wire it into CI

This GitHub Actions template runs two jobs. The first blocks merges and fails on flakes. The second keeps quarantined tests running without letting their failures block, but still fails when the policy is broken. Adapt the install and run steps to your project; the flags and the script are the parts that matter.

name: e2e
 
on: [push, pull_request]
 
jobs:
  tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - run: npm ci
      - run: npx playwright install --with-deps
      # Blocking: quarantined tests are excluded, and a pass on retry fails the run.
      - run: npx playwright test --grep-invert @quarantine --fail-on-flaky-tests
      - if: always()
        run: node flake-report.mjs results/results.json
 
  quarantine:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - run: npm ci
      - run: npx playwright install --with-deps
      # Non-blocking: quarantined tests still run, so their failures stay visible.
      - run: npx playwright test --grep @quarantine --pass-with-no-tests
        continue-on-error: true
      # Blocking: a missing owner, missing issue, or expired quarantine fails the job.
      - if: always()
        run: node flake-report.mjs results/results.json

continue-on-error is set on the test step only, not on the job. If it were on the job, an expired quarantine would never turn anything red.

--pass-with-no-tests covers the good days when nothing is quarantined. Without it, Playwright exits with "No tests found" and the step shows as failed. We checked that the JSON report is still written in both cases, so the policy step always has a file to read.

A quarantine policy you can paste

Put this in your repository's testing guide and adjust the numbers:

# Flaky test quarantine policy
 
1. **Detect.** CI fails any run where a test passes only on retry (`--fail-on-flaky-tests`).
2. **Decide within one working day.** Fix the test, or quarantine it. Re-running the pipeline until it turns green is not an option.
3. **Quarantine with three facts.** Tag the test `@quarantine` and annotate it with `owner=`, `issue=`, and `expires=`, at most 14 days out.
4. **Keep it running.** Quarantined tests run on every pull request in a non-blocking job, so their failures stay visible.
5. **Expire loudly.** CI fails when a quarantine annotation is missing a field or has expired.
6. **Prove the fix.** A test leaves quarantine after the fix passes a repeat run under normal parallelism, for example `--repeat-each=20`.
7. **Delete, don't hoard.** If a test reaches a second expiry without a fix, delete it and record the lost coverage in the issue.
8. **Review weekly.** Track how many tests are quarantined and how old the oldest one is. Both numbers should go down.

Conclusion

Retries are a reasonable safety net, but a flaky test that passes on retry leaves no trace in Playwright's JUnit report. Add the JSON reporter, fail runs on flakes, and record traces on the first retry. When a test flakes, measure how often before changing anything, confirm your experiments actually ran tests, and change one factor at a time. When a test can't be fixed today, quarantine it with an owner, an issue, and an expiry date that CI enforces.

If you need flake history across many runs rather than per run, a dedicated tool such as DeFlaky compares results across repeated runs.

Sources and further reading

Tools mentioned

PlaywrightUI AutomationOpen source
GitHub ActionsCI/CDFree plan

Links go to each tool’s official site. How we choose and link tools

Revision history

Updated source links that had moved: the pages still exist, at new addresses.

Spotted a mistake? Report it — corrections land here.

Written and reviewed by

QA Vibes Editorial

Articles are written and reviewed by practicing QA and automation engineers. Every article lists its sources and shows when it was last updated.

Use your own numbers

Flaky test rerun calculator

Enter how often a test failed, and see how many clean runs it takes to trust a fix.