Introduction
A flaky test is one that fails and then passes without any code change. With retries enabled, Playwright runs a failed test again, and if the second attempt passes, the run can still exit green. That is useful: one slow network call shouldn't block a release. It's also dangerous: the pipeline stays green while a real race condition sits in the product or the test.
This guide covers three things, all checked on Playwright 1.63.0:
- What a retried flaky test actually looks like in each reporter, including one that records nothing at all.
- How we diagnosed a real flaky test on this site, with the numbers.
- A quarantine policy with owners and expiry dates, enforced by a short script in CI.
What a retried flake looks like in each reporter
We built a small project with three tests and retries: 1: one stable test, one that fails on its first attempt and passes on the retry, and one quarantined test we'll use later. The flaky test is deterministic, so every run shows the same result:
import { test, expect } from "@playwright/test";
// Deterministically flaky: fails on the first attempt, passes on the retry.
test("checkout button navigates", async ({ browserName }, testInfo) => {
expect(testInfo.retry, "simulated flake on the first attempt").toBeGreaterThan(0);
});The config writes three reports at once:
import { defineConfig } from "@playwright/test";
export default defineConfig({
testDir: "./tests",
retries: 1,
reporter: [
["list"],
["json", { outputFile: "results/results.json" }],
["junit", { outputFile: "results/results.xml" }],
],
});Running the stable and flaky tests gave this:
| Where you look | What it showed |
|---|---|
| Exit code | 0: the run passed |
| List reporter (terminal) | 1 flaky, 1 passed |
| JUnit XML | tests="2" failures="0", no retry, no flaky marker |
| JSON report | "flaky": 1, test status "flaky", attempts failed then passed |
Exit code with --fail-on-flaky-tests |
1: the run failed |
The JUnit report records a clean pass
This is the JUnit XML from that run, unedited:
<testsuites id="" name="" tests="2" failures="0" skipped="0" errors="0" time="0.675869">
<testsuite name="flaky.spec.ts" timestamp="2026-09-13T19:50:47.230Z" hostname="" tests="1" failures="0" skipped="0" time="0.008" errors="0">
<testcase name="checkout button navigates" classname="flaky.spec.ts" time="0.008">
<system-out>
<![CDATA[
[[ATTACHMENT|..\..\..\test-results\flaky-checkout-button-navigates\error-context.md]]
]]>
</system-out>
</testcase>
</testsuite>
<testsuite name="stable.spec.ts" timestamp="2026-09-13T19:50:47.230Z" hostname="" tests="1" failures="0" skipped="0" time="0.004" errors="0">
<testcase name="cart total adds up" classname="stable.spec.ts" time="0.004">
</testcase>
</testsuite>
</testsuites>The flaky test has no <failure> element, no retry count, and no flaky marker. The only hint is an attachment path to the failed attempt's error context. A dashboard or test management tool that reads only this file sees two passing tests.
This is a known limitation, not a bug. A 2025 request to add an opt-in flaky marker to the JUnit reporter was closed with a "collecting feedback" label. Maintainers have preferred to keep that reporter standard JUnit.
The JSON report records the truth
Trimmed to the relevant fields, the same run's JSON report looks like this:
{
"stats": { "expected": 1, "skipped": 0, "unexpected": 0, "flaky": 1 },
"tests": [
{
"title": "checkout button navigates",
"status": "flaky",
"attempts": [
{ "retry": 0, "status": "failed" },
{ "retry": 1, "status": "passed" }
]
}
]
}In the real report, tests are nested under suites[].specs[].tests[], and each attempt lives in results[]. The script later in this guide walks that structure.
Step 1: Make flakes visible
Keep retries if you need them, but stop letting them hide flakes:
- Add the JSON reporter next to whatever reporter feeds your dashboard. The terminal output says a test was flaky, but JSON is the report file a script can read that still says so.
- Fail the run on flakes with
--fail-on-flaky-tests. The retry still tells you whether the failure was persistent, but the pipeline goes red either way. - Record traces on the first retry with
trace: "on-first-retry"in theuseblock of your config. Every flake then leaves a trace you can open in the trace viewer.
npx playwright test --fail-on-flaky-testsStep 2: Reproduce before you fix
This site's own test suite had a flaky test: the smoke test that clicks through every link in the main navigation. Here is how we found the cause, including two mistakes along the way.
Measure a rate, not an anecdote. A single failure proves nothing. We ran the test repeatedly with the same parallelism as a normal run:
npx playwright test tests/e2e/smoke.spec.ts --grep "main navigation" --repeat-each=20 --workers=6 --output=.flake-evidenceIt failed 3 times in 15. The --output folder matters: every new run clears the default test-results folder, and our first round of evidence was deleted exactly that way.
Mistake 1: a fix that only moved the failure. The test waited for a heading to be visible before checking which link was active. We assumed a timing race and made it wait for the URL instead. Failures dropped but didn't stop. The traces showed why: the click completed, the page data request returned 200 in 16 to 191 ms, and the page still never changed. Nothing was slow. The navigation simply never committed.
Mistake 2: a measurement that ran zero tests. To isolate the cause, we wrote a small experiment and filtered it with --grep. The filter matched nothing, Playwright reported "No tests found", and our script counted that as zero failures. Always check the number of tests a run executed before trusting its failure count.
Change one factor at a time. The experiment clicked a second link while the first navigation was still finishing:
| Configuration | Runs | Second click lost |
|---|---|---|
| Next.js 16.0.3, 6 workers, header links | 40 | 8 |
| Next.js 16.0.3, 6 workers, footer links with no click handler | 40 | 5 |
| Next.js 16.0.3, 1 worker | 80 | 0 |
| Next.js 16.3.5, 6 workers | 80 | 0 |
Footer links failed too, which cleared our own menu code. With a single worker, nothing failed, so the behavior needed CPU contention to appear. Upgrading the framework removed it under the same load, and the smoke test then passed 40 out of 40 runs on desktop and mobile. No release note names this fix, so we only claim what we measured on our machine. It's also why we now pin exact framework versions.
What we did not do: add a longer wait, add a retry, or mark the test fixme. Each of those would have turned a measurable problem into an invisible one. Our CI retries failed tests once, so there this flake would usually have passed on the retry and shown up as green.
Many flakes have more ordinary causes: fixed waits, missing await, or assertions that don't retry. Our Playwright test reviewer checks a pasted test for those patterns.
Step 3: Quarantine with an owner and an expiry date
Sometimes a flaky test can't be fixed today. Quarantine lets the rest of the suite keep blocking merges while that one test stops blocking. Done badly, quarantine becomes a permanent skip list. Done well, every quarantined test still runs, has an owner, links to an issue, and expires.
Tag and annotate the test
Playwright tags and annotations carry that information in the test itself:
import { test, expect } from "@playwright/test";
test(
"search suggestions appear",
{
tag: "@quarantine",
annotation: {
type: "quarantine",
description: "owner=@qa-lead issue=https://github.com/example/shop/issues/482 expires=2026-10-15",
},
},
async () => {
expect("no suggestions").toBe("3 suggestions");
},
);--grep-invert @quarantine runs everything except quarantined tests, and --grep @quarantine runs only them. Both tags and annotations appear in the JSON report. One detail cost us a few minutes: the report stores the tag as quarantine, without the @, so a script that looks for "@quarantine" never matches.
Enforce the policy with the JSON report
This script reads the JSON report. It prints a warning for each flaky test, notes quarantined tests that pass again, and fails if a quarantine is missing its owner, issue, or expiry date, or has expired. On GitHub Actions it also adds a table of flaky tests to the job summary.
// flake-report.mjs: reads Playwright's JSON report, lists flaky tests, and enforces a
// quarantine policy. Every @quarantine test needs an owner, an issue, and an expiry date.
// Usage: node flake-report.mjs results/results.json
import fs from "node:fs";
const report = JSON.parse(fs.readFileSync(process.argv[2] ?? "results/results.json", "utf8"));
const today = new Date().toISOString().slice(0, 10);
const flaky = new Map();
const problems = new Set();
const recovered = new Set();
function visit(suite) {
for (const spec of suite.specs ?? []) {
for (const test of spec.tests) {
const name = `${spec.file} › ${spec.title}`;
if (test.status === "flaky") {
flaky.set(name, test.results.map((r) => r.status).join(" → "));
}
// Playwright's JSON report stores tags without the leading "@".
if (!spec.tags.includes("quarantine")) continue;
const note = test.annotations.find((a) => a.type === "quarantine")?.description ?? "";
const owner = note.match(/owner=(\S+)/)?.[1];
const issue = note.match(/issue=(\S+)/)?.[1];
const expires = note.match(/expires=(\d{4}-\d{2}-\d{2})/)?.[1];
if (!owner || !issue || !expires) {
problems.add(`${name}: quarantine annotation needs owner=, issue=, and expires=YYYY-MM-DD`);
} else if (expires < today) {
problems.add(`${name}: quarantine expired on ${expires} (owner ${owner}, ${issue})`);
}
if (test.status === "expected") recovered.add(name);
}
}
for (const child of suite.suites ?? []) visit(child);
}
report.suites.forEach(visit);
for (const [name, attempts] of flaky) console.log(`::warning title=Flaky test::${name} (${attempts})`);
for (const name of recovered) console.log(`::notice title=Quarantined test passed::${name}. If it keeps passing, take it out of quarantine.`);
for (const problem of problems) console.log(`::error title=Quarantine policy::${problem}`);
if (process.env.GITHUB_STEP_SUMMARY) {
const rows = [...flaky].map(([name, attempts]) => `| ${name} | ${attempts} |`).join("\n");
const table = flaky.size ? `| Test | Attempts |\n| --- | --- |\n${rows}\n` : "None this run.\n";
fs.appendFileSync(process.env.GITHUB_STEP_SUMMARY, `### Flaky tests: ${flaky.size}\n\n${table}`);
}
console.log(`${flaky.size} flaky, ${recovered.size} recovered in quarantine, ${problems.size} policy problem(s)`);
process.exit(problems.size > 0 ? 1 : 0);We ran it against the real reports from our example project, plus copies edited to cover each failure case:
flaky run ::warning title=Flaky test::flaky.spec.ts › checkout button navigates (failed → passed)
1 flaky, 0 recovered in quarantine, 0 policy problem(s) exit 0
valid quarantine 0 flaky, 0 recovered in quarantine, 0 policy problem(s) exit 0
expired ::error title=Quarantine policy::quarantine.spec.ts › search suggestions appear: quarantine expired on 2026-09-01 (owner @qa-lead, https://github.com/example/shop/issues/482)
exit 1
missing owner ::error title=Quarantine policy::quarantine.spec.ts › search suggestions appear: quarantine annotation needs owner=, issue=, and expires=YYYY-MM-DD
exit 1
passing again ::notice title=Quarantined test passed::quarantine.spec.ts › search suggestions appear. If it keeps passing, take it out of quarantine.
exit 0On GitHub Actions, the ::warning, ::notice, and ::error lines become annotations on the run.
Wire it into CI
This GitHub Actions template runs two jobs. The first blocks merges and fails on flakes. The second keeps quarantined tests running without letting their failures block, but still fails when the policy is broken. Adapt the install and run steps to your project; the flags and the script are the parts that matter.
name: e2e
on: [push, pull_request]
jobs:
tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: npx playwright install --with-deps
# Blocking: quarantined tests are excluded, and a pass on retry fails the run.
- run: npx playwright test --grep-invert @quarantine --fail-on-flaky-tests
- if: always()
run: node flake-report.mjs results/results.json
quarantine:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: npx playwright install --with-deps
# Non-blocking: quarantined tests still run, so their failures stay visible.
- run: npx playwright test --grep @quarantine --pass-with-no-tests
continue-on-error: true
# Blocking: a missing owner, missing issue, or expired quarantine fails the job.
- if: always()
run: node flake-report.mjs results/results.jsoncontinue-on-error is set on the test step only, not on the job. If it were on the job, an expired quarantine would never turn anything red.
--pass-with-no-tests covers the good days when nothing is quarantined. Without it, Playwright exits with "No tests found" and the step shows as failed. We checked that the JSON report is still written in both cases, so the policy step always has a file to read.
A quarantine policy you can paste
Put this in your repository's testing guide and adjust the numbers:
# Flaky test quarantine policy
1. **Detect.** CI fails any run where a test passes only on retry (`--fail-on-flaky-tests`).
2. **Decide within one working day.** Fix the test, or quarantine it. Re-running the pipeline until it turns green is not an option.
3. **Quarantine with three facts.** Tag the test `@quarantine` and annotate it with `owner=`, `issue=`, and `expires=`, at most 14 days out.
4. **Keep it running.** Quarantined tests run on every pull request in a non-blocking job, so their failures stay visible.
5. **Expire loudly.** CI fails when a quarantine annotation is missing a field or has expired.
6. **Prove the fix.** A test leaves quarantine after the fix passes a repeat run under normal parallelism, for example `--repeat-each=20`.
7. **Delete, don't hoard.** If a test reaches a second expiry without a fix, delete it and record the lost coverage in the issue.
8. **Review weekly.** Track how many tests are quarantined and how old the oldest one is. Both numbers should go down.Conclusion
Retries are a reasonable safety net, but a flaky test that passes on retry leaves no trace in Playwright's JUnit report. Add the JSON reporter, fail runs on flakes, and record traces on the first retry. When a test flakes, measure how often before changing anything, confirm your experiments actually ran tests, and change one factor at a time. When a test can't be fixed today, quarantine it with an owner, an issue, and an expiry date that CI enforces.
If you need flake history across many runs rather than per run, a dedicated tool such as DeFlaky compares results across repeated runs.
Sources and further reading
- Playwright: retries
- Playwright: reporters
- Playwright: annotations and tags
- Playwright: command line (
--fail-on-flaky-tests,--repeat-each) - Playwright: trace viewer
- Playwright issue #35592: flaky tests in the JUnit reporter
- GitHub Actions: workflow syntax (
continue-on-error) - GitHub Actions: workflow commands and job summaries
- DeFlaky
- Three bugs we caught building this site
- Try it: Playwright test reviewer