Introduction
Four metrics come up when a team asks how many bugs got past its testing: defect leakage, escaped defect rate, defect removal efficiency (DRE), and defect detection percentage (DDP). They sound like four measurements. They're one fraction, looked at from two sides, with a few ways to get it wrong.
This guide gives each formula with where it comes from, shows the denominator mistake that published guides make, explains why the time window matters, puts the widely quoted 85% average in context, and ends with the question the formulas don't answer: whether a change between two releases is real.
The numbers in the examples are illustrative, not measurements from a real project. The calculations were run with the same code as our escaped defect calculator.
One fraction, two sides
Every one of these metrics starts from two counts for a release:
- B: defects found before release, by reviews and testing.
- A: defects found after release, by users or production monitoring.
Then:
- Share caught before release = B ÷ (A + B)
- Share that got through = A ÷ (A + B) = 1 − (share caught)
| Name | Which side | Formula | Where the definition comes from |
|---|---|---|---|
| Defect removal efficiency (DRE) | Caught | B ÷ (A + B) | Capers Jones: the percentage of defects "found and repaired prior to release", with customer-reported defects counted for 90 days |
| Defect detection percentage (DDP) | Caught | Found by a test level ÷ (found by it + found afterwards) | ISTQB Glossary v4.8.1; the same fraction, for any test level, not just the whole release |
| Escaped defect rate | Got through | A ÷ (A + B) | Common in agile teams; we found no standards-body definition |
| Defect leakage | Got through | A ÷ (A + B), or A ÷ B in some guides | Vendor and blog guides; see the next section |
So a release where testing found 90 defects and users reported 10 has a DRE of 90% and an escape rate of 10%. That's one measurement: report whichever side your audience reads more naturally, and don't present both as if they were independent evidence.
DDP is the most general form. The ISTQB Glossary defines it for a test level: the defects found by, say, system testing, divided by those found by system testing plus everything found after it (acceptance testing, production). That's useful when you want to know which stage is letting defects through, not just whether the release did.
Check the denominator
Some guides divide escapes by all defects found: A ÷ (A + B). Others divide by defects found before release: A ÷ B. Both are called defect leakage, and they're different numbers:
| Found before (B) | Found after (A) | A ÷ (A + B) | A ÷ B |
|---|---|---|---|
| 90 | 10 | 10.0% | 11.1% |
| 50 | 10 | 16.7% | 20.0% |
| 50 | 50 | 50.0% | 100.0% |
| 20 | 30 | 60.0% | 150.0% |
When few defects escape, the two are close. When many do, A ÷ B runs past 100%, which is a sign it isn't the share of anything.
The confusion is real in published guides. Of three well-ranked articles on defect leakage we read on 21 September 2026, Testsigma and Ranorex both divide by the total, and their examples match. testRigor writes the formula as defects after release divided by defects before release, but its worked example divides 10 by 90 + 10 and gets 10%. The formula as written would give 11.1%.
Whichever you use, write the formula next to the number every time you report it. A dashboard that says "leakage: 20%" without it can't be compared with anyone else's, or with itself after someone changes the query.
Fix the window for "after"
A is never final: users keep finding old defects for as long as the release is in use. So an escape rate calculated a week after release will be lower than the same release's rate a month later, even though nothing changed.
Capers Jones's description of DRE fixes the window: keep records of all defects found during development, and "after a fixed period of 90 days, add customer-reported defects". Pick a window, write it down, and only compare releases that have all had the full window. A release shipped last week can't be compared with one that has had three months.
Also decide what counts as a defect before you start: duplicates, "works as designed", and feature requests filed as bugs all change A and B. The rule matters less than applying the same rule to every release.
Is there an industry standard?
People search for an "industry standard" or "benchmark" for these metrics. The best-known published figure we found comes from Capers Jones's paper Software Defect Removal Efficiency, which says "the current U.S. average in 2011 is only about 85%" and that "best in class projects can top 99%".
Three things to keep in mind before you use it:
- It's an estimate in one consultant's paper, for 2011, not a standard that any standards body publishes.
- It's DRE with a 90-day window and his definition of a defect. Your count will differ.
- It includes defects removed by reviews and inspections as well as testing, which is one of the paper's main points: it says high removal efficiency "cannot be achieved using testing alone".
We found no current, independent benchmark for escape rate or leakage. The most useful comparison is your own trend, measured the same way each release.
When is a change between releases real?
Here's an illustrative series of four releases, run through the calculator's code:
| Release | Found before (B) | Found after (A) | Escape rate | DRE | 95% range for the escape rate |
|---|---|---|---|---|---|
| 4.9 | 31 | 4 | 11.4% | 88.6% | 4.5% to 26.0% |
| 4.10 | 24 | 6 | 20.0% | 80.0% | 9.5% to 37.3% |
| 4.11 | 28 | 3 | 9.7% | 90.3% | 3.3% to 24.9% |
| 4.12 | 19 | 9 | 32.1% | 67.9% | 17.9% to 50.7% |
From 4.9 to 4.10 the escape rate nearly doubles, from 11.4% to 20.0%. It would be easy to call that a regression. But the range for the change runs from −9.3 to +27.2 percentage points, and it includes zero: a release that was genuinely no different could easily have produced those counts. With 4 and 6 escapes, you can't tell.
From 4.11 to 4.12 the change is +22.5 points with a range of +1.6 to +42.0. The whole range is above zero, so that one is worth a look at which stage stopped catching defects. Against the three earlier releases pooled, 4.12 is +18.6 points, range +2.2 to +37.9: the same conclusion.
The ranges are 95% Wilson intervals for each release, and Newcombe's method for the difference between two releases, which stays within possible values even when a release has no escapes at all. The calculator shows the same numbers for your own counts, and the go or no-go guide covers what to do with them before a release.
Defect density is a different measure
"Defect density" often appears in the same lists, but it answers a different question. The ISTQB Glossary defines it as the number of defects per unit size of a work product: defects per thousand lines, per function point, or per story point. It says how many defects there were for the size of the thing, not how many got past testing. Don't combine it with the metrics above into one score; report it separately, with its unit.
Conclusion
Defect removal efficiency, defect detection percentage, escaped defect rate, and defect leakage all come from two counts: defects found before release and defects found after it. Write the formula and the denominator next to every number, fix how long "after" lasts, treat the 85% figure as one 2011 estimate rather than a target, and check whether a change is bigger than the noise before you act on it.
Sources and further reading
- Capers Jones, Software Defect Removal Efficiency (PDF)
- ISTQB Glossary v4.8.1: defect detection percentage
- ISTQB Glossary v4.8.1: defect density
- Testsigma: What is defect leakage in software testing?
- Ranorex: What is defect leakage in QA?
- testRigor: What is defect leakage in software testing?
- Newcombe, Interval estimation for the difference between independent proportions (Statistics in Medicine, 1998)