Free tool
Flaky test rerun calculator
How flaky is this test, and how many passing runs prove the fix worked? The defaults are real numbers from a flaky test on this site.
Step 1
Measure the flake rate
Run the unchanged test many times, with the same parallelism as CI, and count the failures.
Failure rate
20% (3 of 15 runs). The true rate is probably between 7.0% and 45% (95% range).
Runs to prove a fix (95% confidence)
After your fix, the test must pass 14 runs in a row. If the real rate is at the low end (7.0%), you need 41.
Command
npx playwright test path/to/test.spec.ts --repeat-each=41Step 2
Check a fix
Enter how many runs in a row passed after the fix, again with CI's parallelism.
What 40 clean runs rule out (95%)
The failure rate is now at most 7.2%. Anything higher would very likely have failed at least once.
If the fix didn't work
A test still failing 20% of the time would pass all 40 runs by luck with a 0.013% chance. Your clean runs make that unlikely.
How the math works
Failure rate and range. The rate is failures divided by runs. The 95% range is a Wilson score interval, which stays sensible for small counts: 3 failures in 15 runs could easily come from a test that really fails 7% or 45% of the time.
Runs to prove a fix. If the test still failed at rate p, the chance it passes n runs in a row by luck is (1 − p)n. The calculator finds the smallest n that makes that chance lower than 5% (for 95% confidence). A test that failed 20% of the time needs 14 clean runs; one that failed 5% of the time needs 59.
What clean runs rule out. After n clean runs, any failure rate above 1 − (1 − confidence)1/n would very likely have shown at least one failure. At 95% this is close to the rule of three: about 3 ÷ n. Forty clean runs rule out rates above about 7%.
When the numbers can mislead you
- The math assumes every run is independent and the rate doesn't change. Flakes caused by CPU load or shared data break that: our own flaky test failed only with six parallel workers and never with one. Rerun with the same parallelism as CI, for example
--repeat-eachwith your CI's--workersvalue. - Clean runs on your laptop say little about CI. Measure where the test actually fails.
- Rare flakes need many runs. A test that fails 1% of the time needs about 300 clean runs to prove a fix at 95% confidence, so fixing the cause usually beats rerunning.
The calculator runs in your browser and sends nothing. For how we measured, diagnosed, and fixed the flaky test behind the defaults, read Flaky tests hide behind retries.