Introduction
The test pyramid is a picture of how a test suite should be shaped: many small, fast tests at the bottom, fewer broader tests in the middle, and a thin layer of slow end-to-end tests at the top. Mike Cohn introduced it in Succeeding with Agile (2009), and Martin Fowler's short 2012 write-up made it the default mental model.
The picture is simple. What people do with it is not. Teams quote percentages that nobody published, argue about layer names, and treat the shape as a target instead of a warning sign. This guide sticks to what the sources say, shows one example per layer, and then looks at a real suite — this site's — that doesn't match the textbook shape.
The three layers
The layers are defined by scope, not by tool. A Playwright test can be an API test; a Jest test can start a database.
- Unit (narrow scope). One function or class, no network, no browser. Runs in milliseconds. When it fails, the cause is in a few lines of code.
- Integration or service (medium scope). Two or more components talking for real: an HTTP handler and its database, or a client and a contract. Runs in tens or hundreds of milliseconds.
- End-to-end (large scope). The deployed system through its real interface, usually a browser. Runs in seconds. When it fails, the cause can be anywhere.
The trade-off is the whole argument: as scope grows, each test gives more confidence that the system works, but it runs slower, fails for more reasons, and points less precisely at the bug.
One test per layer
The same checkout feature, tested at each layer.
A unit test for the tax rule. 100.05 × 8.25% is 8.254125, and the rule is to round up to the next cent:
// src/cart/tax.test.ts (Jest)
import { calcTax } from "./tax";
test("rounds tax up to the next cent", () => {
expect(calcTax(100.05, 0.0825)).toBe(8.26);
});A service test for the orders endpoint, using Playwright's request fixture, so no browser starts:
// tests/api/orders.spec.ts
import { test, expect } from "@playwright/test";
test("POST /api/orders returns the created order", async ({ request }) => {
const response = await request.post("/api/orders", {
data: { sku: "TSHIRT-M", quantity: 1 },
});
expect(response.status()).toBe(201);
expect(await response.json()).toMatchObject({ sku: "TSHIRT-M", quantity: 1 });
});An end-to-end test for the journey a customer actually takes:
// tests/e2e/checkout.spec.ts
import { test, expect } from "@playwright/test";
test("guest can place an order", async ({ page }) => {
await page.goto("/checkout");
await page.getByLabel("Card number").fill("4242 4242 4242 4242");
await page.getByRole("button", { name: "Place order" }).click();
await expect(page.getByRole("heading", { name: "Order placed" })).toBeVisible();
});Notice what each one does not test. The unit test doesn't know there is an API. The API test doesn't care how the rounding works, only that an order comes back. The browser test doesn't check the tax amount at all — if it did, a rounding bug would fail a slow test with an unhelpful message instead of a fast one that names the function.
An earlier version of this article used
page.fill("#card", …)andpage.click("text=Submit")in the browser example. Those selector shorthands are discouraged in Playwright's API docs, and our own Playwright test reviewer didn't flag them. We added the check, and every Playwright example on this site now has to pass the reviewer in CI.
What the published numbers say
Most percentages you'll see in blog posts ("60/30/10 for startups") have no source. Two figures do, and both come from Google:
| Source | Unit | Integration | End-to-end |
|---|---|---|---|
| Google Testing Blog, "Just Say No to More End-to-End Tests" (2015) | 70% | 20% | 10% |
| Software Engineering at Google, chapter 11 (2020) | 80% | 15% | 5% |
The blog post offers 70/20/10 as a first guess and says the exact mix will differ by team. The book describes 80/15/5 as what Google tends to aim for. Neither is a standard, and both come from an organization whose code is mostly backend services with a lot of business logic — which is exactly where unit tests pay off most.
The useful part of both sources is the reasoning, not the numbers: prefer the smallest test that can catch the bug.
Our own suite doesn't match, on purpose
When this article was updated in September 2026, this site's test suite had:
- 789 unit-level tests. 230 of them check the articles themselves: every snippet must parse, every Playwright and Cypress example must pass our own test reviewers, and every article must keep a fixed structure. The rest cover the rules inside the bug report grader and the Playwright test reviewer, the flaky test calculator's math, the Cypress reviewer's rules, the pairwise generator and its release workflow, the shared core behind the two test reviewers, the downloadable templates, the site search's matching rules and synonym list, the roadmap data, and the test design tools, whose definitions of boundary values, decision tables and state transition coverage are checked against the ISTQB syllabus's own examples and against hundreds of generated inputs, and the practice questions, whose answers are recomputed from each question's wording, with the statement and branch coverage minimums checked against a brute-force search over every path.
- 418 browser tests on desktop Chrome and 81 on a mobile viewport. Seven of the desktop ones only run against the deployed site, and skip locally. Smoke checks for every page, a check that every article fits the screen, navigation, SEO tags, accessibility scans, the tools' user interfaces, the practice lab's reference checks, which confirm that every seeded bug is still there, a re-run of every test the diagnose-a-result case files and the AI test challenge show, and every test case the state transition tool generates for the practice shop, run against the shop.
That's roughly 60% unit and 40% browser, with almost nothing in between. By the numbers above it's an hourglass.
It's still the right shape for this project. A content site has little business logic: the product is rendered HTML, so many real bugs — a broken link color, a page that scrolls sideways on mobile, a navigation click that got lost — can only be seen in a browser. Where logic does exist, in the tools, it's unit-tested, and the browser tests only confirm the wiring. The three bugs we wrote up while building this site split the same way: link colors silently overridden by CSS could only be seen in a browser, while the grader approving its own blank template and a placeholder value colliding with real text were logic bugs, fixed with unit-level regression tests — the layer the pyramid says they belong in.
The point isn't that ratios don't matter. It's that the right question is "what kind of bug does this project have, and what's the cheapest test that catches it?" For a pricing engine, that's a unit test almost every time. For a marketing site, it's usually a browser.
Shapes that go wrong
Software Engineering at Google names two antipatterns:
- Ice cream cone: many end-to-end tests, few integration or unit tests. Often the result of a manual test plan being automated one script at a time. The suite is slow, flaky, and every failure needs investigation.
- Hourglass: many unit and end-to-end tests but few integration tests. Failures that a medium-sized test would have caught quickly surface instead as slow, hard-to-diagnose browser failures.
The testing trophy, proposed by Kent C. Dodds for front-end JavaScript, is a deliberate variation rather than an antipattern: static analysis (types, linting) as the base, and the largest share of effort on integration tests, because in UI code the interactions between components are where the bugs are.
Choosing the layer for a new test
- Write it at the lowest layer that can fail for the bug you care about. If the bug is in a calculation, a browser test is the wrong tool.
- When a bug escapes to production, ask which layer should have caught it. Add the test there, not automatically at the top.
- Keep end-to-end tests for journeys, not rules. One test that a guest can check out; unit tests for every tax, discount, and currency rule.
- Measure the suite by time and failure causes, not only counts. If the browser layer takes 80% of the pipeline time and produces most of the flaky failures, that's the shape problem, whatever the percentages say.
- Revisit when the architecture changes. Splitting a monolith into services moves risk into the integration layer; contract tests belong there.
Conclusion
The pyramid is a reminder that test scope has a cost. Use the Google figures as a sanity check for logic-heavy code, not as a quota, and be ready to explain in one sentence why your suite has the shape it has. If you can't, it's probably an ice cream cone that grew by accident.