Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

Free tool

Escaped defect calculator

How many defects reached your users, and is this release actually different from the last one? Enter the counts and find out whether the change you are looking at is real.

Your releases, oldest first

ReleaseFound beforeEscapedRate (95% range)Remove
11% (4.5%–26%)
20% (9.5%–37%)
9.7% (3.3%–25%)
32% (18%–51%)

Everything is calculated in your browser.

What counts as what

Found before is every defect anyone recorded before the release reached users, at any stage. Escaped is every defect found in production afterwards, whoever reported it. Use one rule and keep it: the number is only comparable with your own previous releases, never with another team’s.

Escaped defect rate, all releases pooled

18%

95% range 12%–25% · 22 escaped of 124 found

Defect removal efficiency is 82% — the same measurement seen from the other side, not a second piece of evidence.

4.12 vs 4.11

+22 points

the change could be anywhere from +1.6 points to +42 points

More defects escaped, and there are enough of them to say so: the whole range for the change is above zero. This is worth a look at which gate stopped catching them.

4.12 vs the 3 releases before it

+19 points

the change could be anywhere from +2.2 points to +38 points

More defects escaped, and there are enough of them to say so: the whole range for the change is above zero. This is worth a look at which gate stopped catching them.

Comparing the latest release with its whole history uses more of your data than comparing it with the one before, so when the two disagree, trust the longer one. Both are ranges rather than verdicts: a proportion measured on a few dozen defects moves around a great deal on its own.

What it calculates

The escaped defect rate is the defects found in production divided by every defect found, before and after. A release where 4 of 35 escaped has a rate of 11%. Defect removal efficiency is the other 89%; it is the same measurement, not a second one, so moving both onto a dashboard makes a team look twice as instrumented as it is. Defect leakage and defect detection percentage are the same fraction again; how the names, formulas, and denominators line up is in Defect leakage, escape rate, DRE and DDP.

The 95% range is a Wilson score interval, the same one behind the flaky test calculator. It is the part most escaped-defect dashboards leave out, and it is usually the part that matters: 4 escapes out of 35 is consistent with a true rate anywhere from about 4% to 26%.

Why “we halved it” is usually not true

Four escapes in one release and two in the next looks like a 50% improvement, and on those counts it is indistinguishable from a release that was no different at all. So the comparison here is not the two percentages side by side; it is a range for the change between them, calculated with Newcombe’s method from the two Wilson intervals. When that range includes zero, the honest reading is that you cannot tell the releases apart yet.

Newcombe’s method rather than the textbook normal approximation, because that one starts producing impossible answers — a change of minus 140 percentage points — exactly when a release catches everything and nothing escapes, which is the case a team most wants to look at.

The calculator also compares the latest release with all the earlier ones pooled. That uses more of your data than the release immediately before, so when the two comparisons disagree, the longer one is the better guide.

How to keep the number honest

  • Compare with your own history, not a benchmark. What counts as a defect varies so much between teams that published figures are not comparable. Your previous releases were counted by the same people under the same rules.
  • Never make it a target. The rate improves immediately if fewer production defects get recorded, and that is the easiest way to move it. Use it as a trigger for a conversation, not a goal on anyone’s objectives.
  • Give escapes time to arrive. A release measured the morning after looks excellent. Pick a window — a fortnight, a month — and use the same one every time.
  • One rule for what counts. If duplicates, support questions, or known issues are counted one release and not the next, the trend is measuring your process for recording defects.

The example numbers are illustrative, not measurements from this site: we have not shipped enough releases with recorded defect counts to publish a real series, and inventing one would be worse than saying so.

What to do with a rise is in go or no-go — every Major escape deserves one question about which gate should have caught it. Where the escapes are first visible is usually the support queue, covered in reading quality from the support queue.