Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

Understanding the Test Pyramid: What the Published Numbers Actually Say

Where the test pyramid comes from, the only widely cited ratios and what their authors said about them, the ice cream cone and hourglass antipatterns, and a rule for choosing the layer of each new test.

QA Vibes EditorialPublished Updated 7 minRevision history ↓

Key takeaways

  • The widely cited ratios come from Google: 70/20/10 as a first guess and 80/15/5 as a target, for logic-heavy code.
  • Write each test at the lowest layer that can catch the bug.
  • Ice cream cone and hourglass suites are slow and hard to debug.
  • This site's suite is browser-heavy on purpose, because its product is rendered pages.
Contents (9 sections)

Introduction

The test pyramid is a picture of how a test suite should be shaped: many small, fast tests at the bottom, fewer broader tests in the middle, and a thin layer of slow end-to-end tests at the top. Mike Cohn introduced it in Succeeding with Agile (2009), and Martin Fowler's short 2012 write-up made it the default mental model.

The picture is simple. What people do with it is not. Teams quote percentages that nobody published, argue about layer names, and treat the shape as a target instead of a warning sign. This guide sticks to what the sources say, shows one example per layer, and then looks at a real suite — this site's — that doesn't match the textbook shape.

The three layers

The layers are defined by scope, not by tool. A Playwright test can be an API test; a Jest test can start a database.

  1. Unit (narrow scope). One function or class, no network, no browser. Runs in milliseconds. When it fails, the cause is in a few lines of code.
  2. Integration or service (medium scope). Two or more components talking for real: an HTTP handler and its database, or a client and a contract. Runs in tens or hundreds of milliseconds.
  3. End-to-end (large scope). The deployed system through its real interface, usually a browser. Runs in seconds. When it fails, the cause can be anywhere.

The trade-off is the whole argument: as scope grows, each test gives more confidence that the system works, but it runs slower, fails for more reasons, and points less precisely at the bug.

One test per layer

The same checkout feature, tested at each layer.

A unit test for the tax rule. 100.05 × 8.25% is 8.254125, and the rule is to round up to the next cent:

// src/cart/tax.test.ts (Jest)
import { calcTax } from "./tax";
 
test("rounds tax up to the next cent", () => {
  expect(calcTax(100.05, 0.0825)).toBe(8.26);
});

A service test for the orders endpoint, using Playwright's request fixture, so no browser starts:

// tests/api/orders.spec.ts
import { test, expect } from "@playwright/test";
 
test("POST /api/orders returns the created order", async ({ request }) => {
  const response = await request.post("/api/orders", {
    data: { sku: "TSHIRT-M", quantity: 1 },
  });
  expect(response.status()).toBe(201);
  expect(await response.json()).toMatchObject({ sku: "TSHIRT-M", quantity: 1 });
});

An end-to-end test for the journey a customer actually takes:

// tests/e2e/checkout.spec.ts
import { test, expect } from "@playwright/test";
 
test("guest can place an order", async ({ page }) => {
  await page.goto("/checkout");
  await page.getByLabel("Card number").fill("4242 4242 4242 4242");
  await page.getByRole("button", { name: "Place order" }).click();
  await expect(page.getByRole("heading", { name: "Order placed" })).toBeVisible();
});

Notice what each one does not test. The unit test doesn't know there is an API. The API test doesn't care how the rounding works, only that an order comes back. The browser test doesn't check the tax amount at all — if it did, a rounding bug would fail a slow test with an unhelpful message instead of a fast one that names the function.

An earlier version of this article used page.fill("#card", …) and page.click("text=Submit") in the browser example. Those selector shorthands are discouraged in Playwright's API docs, and our own Playwright test reviewer didn't flag them. We added the check, and every Playwright example on this site now has to pass the reviewer in CI.

What the published numbers say

Most percentages you'll see in blog posts ("60/30/10 for startups") have no source. Two figures do, and both come from Google:

Source Unit Integration End-to-end
Google Testing Blog, "Just Say No to More End-to-End Tests" (2015) 70% 20% 10%
Software Engineering at Google, chapter 11 (2020) 80% 15% 5%

The blog post offers 70/20/10 as a first guess and says the exact mix will differ by team. The book describes 80/15/5 as what Google tends to aim for. Neither is a standard, and both come from an organization whose code is mostly backend services with a lot of business logic — which is exactly where unit tests pay off most.

The useful part of both sources is the reasoning, not the numbers: prefer the smallest test that can catch the bug.

Our own suite doesn't match, on purpose

When this article was updated in September 2026, this site's test suite had:

  • 789 unit-level tests. 230 of them check the articles themselves: every snippet must parse, every Playwright and Cypress example must pass our own test reviewers, and every article must keep a fixed structure. The rest cover the rules inside the bug report grader and the Playwright test reviewer, the flaky test calculator's math, the Cypress reviewer's rules, the pairwise generator and its release workflow, the shared core behind the two test reviewers, the downloadable templates, the site search's matching rules and synonym list, the roadmap data, and the test design tools, whose definitions of boundary values, decision tables and state transition coverage are checked against the ISTQB syllabus's own examples and against hundreds of generated inputs, and the practice questions, whose answers are recomputed from each question's wording, with the statement and branch coverage minimums checked against a brute-force search over every path.
  • 418 browser tests on desktop Chrome and 81 on a mobile viewport. Seven of the desktop ones only run against the deployed site, and skip locally. Smoke checks for every page, a check that every article fits the screen, navigation, SEO tags, accessibility scans, the tools' user interfaces, the practice lab's reference checks, which confirm that every seeded bug is still there, a re-run of every test the diagnose-a-result case files and the AI test challenge show, and every test case the state transition tool generates for the practice shop, run against the shop.

That's roughly 60% unit and 40% browser, with almost nothing in between. By the numbers above it's an hourglass.

It's still the right shape for this project. A content site has little business logic: the product is rendered HTML, so many real bugs — a broken link color, a page that scrolls sideways on mobile, a navigation click that got lost — can only be seen in a browser. Where logic does exist, in the tools, it's unit-tested, and the browser tests only confirm the wiring. The three bugs we wrote up while building this site split the same way: link colors silently overridden by CSS could only be seen in a browser, while the grader approving its own blank template and a placeholder value colliding with real text were logic bugs, fixed with unit-level regression tests — the layer the pyramid says they belong in.

The point isn't that ratios don't matter. It's that the right question is "what kind of bug does this project have, and what's the cheapest test that catches it?" For a pricing engine, that's a unit test almost every time. For a marketing site, it's usually a browser.

Shapes that go wrong

Software Engineering at Google names two antipatterns:

  • Ice cream cone: many end-to-end tests, few integration or unit tests. Often the result of a manual test plan being automated one script at a time. The suite is slow, flaky, and every failure needs investigation.
  • Hourglass: many unit and end-to-end tests but few integration tests. Failures that a medium-sized test would have caught quickly surface instead as slow, hard-to-diagnose browser failures.

The testing trophy, proposed by Kent C. Dodds for front-end JavaScript, is a deliberate variation rather than an antipattern: static analysis (types, linting) as the base, and the largest share of effort on integration tests, because in UI code the interactions between components are where the bugs are.

Choosing the layer for a new test

  1. Write it at the lowest layer that can fail for the bug you care about. If the bug is in a calculation, a browser test is the wrong tool.
  2. When a bug escapes to production, ask which layer should have caught it. Add the test there, not automatically at the top.
  3. Keep end-to-end tests for journeys, not rules. One test that a guest can check out; unit tests for every tax, discount, and currency rule.
  4. Measure the suite by time and failure causes, not only counts. If the browser layer takes 80% of the pipeline time and produces most of the flaky failures, that's the shape problem, whatever the percentages say.
  5. Revisit when the architecture changes. Splitting a monolith into services moves risk into the integration layer; contract tests belong there.

Conclusion

The pyramid is a reminder that test scope has a cost. Use the Google figures as a sanity check for logic-heavy code, not as a quota, and be ready to explain in one sentence why your suite has the shape it has. If you can't, it's probably an ice cream cone that grew by accident.

Sources and further reading

Tools mentioned

JestUI AutomationOpen source
PlaywrightUI AutomationOpen source
PostmanAPI TestingFree plan

Links go to each tool’s official site. How we choose and link tools

Revision history

Updated this site's test counts after the ads.txt, IndexNow, one-page-per-technique practice, Google Analytics consent and AI test challenge checks. The total in the description had fallen behind; it now matches.
Updated this site's test counts (now roughly 60% unit) after the three test design tools, whose browser tests run every generated practice shop case against the shop, and the test design practice questions.
Updated this site's test counts after adding the navigation row check.
Updated this site's test counts after adding the citation checks.
Updated this site's test counts after adding synonyms to site search.
Updated this site's test counts after adding site search and a redirect for a renamed article.
Updated this site's test counts after adding the Azure DevOps article and its checks.
Updated this site's test counts after adding the viewport width sweep and the competency targets.
Updated this site's test counts after adding the page weight budgets.
Updated this site's test counts after adding the code contrast checks.
Updated this site's test counts after adding the competency matrix.
Updated this site's test counts after adding checks that only run against the deployed site.
Updated this site's test counts after adding the Azure Test Plans test case CSV builder.
Updated this site's test counts after adding the Azure Pipelines test results checker.
Updated this site's test counts after adding checks for the security headers.
Updated this site's test counts after adding the page objects versus fixtures article.
Updated this site's test counts after adding the defect metrics article and its check against the calculator's code.
Updated this site's test counts after adding the API test cases article.
Updated this site's test counts after adding the login test cases article.
Updated this site's test counts after adding checks that every article suggests a tool that fits it, or none.
Updated this site's test counts after adding checks for the tool events that decide what gets built next.
Updated this site's test counts (now roughly 65% unit) after the SQL lab, whose unit tests re-run every query in its article.
Updated this site's test counts (now roughly 64% unit) after the product owner tools and the diagnose-a-test-result case files, whose browser tests re-run every test those pages show.
Updated this site's test counts after adding checks for what search engines may index and for where the practice shop's seeded bugs live.
Updated this site's test counts after the tool pages started reading one list, with a check on each page's own metadata.
Updated this site's test counts after the pairwise generator learned to save and open a test matrix as a file.
Updated this site's test counts and the unit share (now roughly 66%) after the pairwise generator's state moved out of its component and became unit-testable.
Updated this site's test counts after adding the accessibility testing article.
Updated this site's test counts after adding the k6 load testing article.
Updated this site's test counts and the unit share (now roughly 64%) after adding value aliases to the pairwise generator.
Updated this site's test counts after the pairwise generator started keeping case IDs through a release update.
Updated this site's test counts after the pairwise generator learned to keep last release's test cases when the model changes.
Updated this site's test counts after adding a coverage check for existing test cases to the pairwise generator, and corrected the total in the description.
Updated this site's test counts after adding the downloadable bug report and test case templates.
Updated this site's test counts after adding known constructions to the pairwise test case generator.
Updated this site's test counts after gating Cypress examples with the Cypress test reviewer.
Updated this site's test counts after adding compare mode to the Cypress test reviewer.
Updated this site's test counts again after adding the practice lab and its reference checks.
Updated this site's test counts after adding new tools and the pairwise testing article.
Rewritten. Removed unsourced ratios, added the figures Google published and this site's own test counts, and replaced discouraged page.fill and page.click selector shorthands in the browser example.
Revised during a site-wide content audit.
First published.

Spotted a mistake? Report it — corrections land here.

Written and reviewed by

QA Vibes Editorial

Articles are written and reviewed by practicing QA and automation engineers. Every article lists its sources and shows when it was last updated.