Introduction
Most pairwise testing guides stop at the count: this many parameters, this many combinations, only this many test cases. They rarely run the cases against anything, so you never see what the smaller set actually catches or misses.
This guide runs them. We built a pairwise model of the checkout on Sauce Demo, a public practice shop with deliberately broken accounts. We generated the cases with our free pairwise test case generator and ran them as a data-driven Playwright test on Chromium, Firefox, and WebKit. Then we ran every combination of account, product, and form input to find out what the pairwise set missed.
Everything below was run on Playwright 1.63.0 on 2026-09-15. The numbers are from those runs, and the code is the code we ran.
Pairwise testing in one paragraph
Pairwise (all-pairs) testing picks a set of test cases in which every pair of values from any two parameters appears at least once. NIST's combinatorial testing project reports that most software failures are triggered by one or two parameters, with fewer by three or more, based on studies such as Kuhn, Wallace, and Gallo (2004). Covering every pair should therefore catch many interaction bugs with far fewer cases than all combinations. "Many" isn't "all", and the rest of this article measures the difference on one real app.
Step 1: Build the model
A model is a list of parameters and their values. For the checkout, we chose the things a tester would reasonably vary:
Account: standard_user, problem_user, performance_glitch_user, error_user, visual_user
Product: Sauce Labs Backpack, Sauce Labs Bike Light, Sauce Labs Bolt T-Shirt, Sauce Labs Fleece Jacket, Sauce Labs Onesie, Test.allTheThings() T-Shirt (Red)
Browser: chromium, firefox, webkit
Viewport: desktop, mobile
First name: filled, ~empty
Last name: filled, ~empty
Postal code: filled, ~emptyThree modeling decisions matter more than the generator:
- Empty fields are negative values. The
~marks an invalid value. The generator puts at most one in each case, so when a case fails validation, you know which empty field caused it. Without the marker, a case with all three fields empty would only ever show the first error, and the other two checks would never run. locked_out_userisn't in the model. Its login fails before the product, browser, or form can make a difference. As a value in the model, it would still need to be paired with every product, viewport, and form value, adding rows that stop at the login page. It gets one separate test instead.- Form values are "filled" or "empty", not real names. Pairwise testing chooses combinations. Choosing the values themselves is still equivalence partitioning and boundary analysis, covered in how to write effective test cases.
Step 2: Generate the cases
Pasting the model into the generator gives 31 test cases covering all 196 valid pairs. All combinations of these values would be 5 × 6 × 3 × 2 × 2 × 2 × 2 = 1,440 checkouts. Three more pairs exist on paper (two empty fields in one case), and the negative-value rule excludes them.
The smallest possible pairwise set here is 30, because every account must meet every product and each case holds only one account-product pair. For three-way coverage, the same model needs 97 cases.
The generator exports the set as a Gherkin Examples table, so it can feed a Cucumber Scenario Outline directly. These are the first rows:
Examples:
| Account | Product | Browser | Viewport | First name | Last name | Postal code |
| standard_user | Sauce Labs Backpack | chromium | desktop | filled | filled | filled |
| problem_user | Sauce Labs Bike Light | firefox | mobile | filled | empty | filled |
| performance_glitch_user | Sauce Labs Bolt T-Shirt | webkit | mobile | empty | filled | filled |
| error_user | Sauce Labs Fleece Jacket | firefox | desktop | filled | filled | empty |
| visual_user | Sauce Labs Onesie | webkit | desktop | filled | empty | filled |
| error_user | Test.allTheThings() T-Shirt (Red) | chromium | mobile | empty | filled | filled |For Playwright, we exported JSON instead and saved it as cases.json. Each case is one object:
[
{
"Account": "standard_user",
"Product": "Sauce Labs Backpack",
"Browser": "chromium",
"Viewport": "desktop",
"First name": "filled",
"Last name": "filled",
"Postal code": "filled"
}
]Step 3: Run the cases as a data-driven Playwright test
The checkout test
One function registers a checkout test for one row of data. Every step has an assertion with a message, so a failure says where the checkout broke, not just that it did:
import { test, expect } from "@playwright/test";
export type Case = {
Account: string;
Product: string;
Browser?: string;
Viewport?: string;
"First name": "filled" | "empty";
"Last name": "filled" | "empty";
"Postal code": "filled" | "empty";
};
const PASSWORD = process.env.SAUCE_PASSWORD ?? "secret_sauce";
const fields = [
{ key: "First name", testId: "firstName", value: "Ada", error: "Error: First Name is required" },
{ key: "Last name", testId: "lastName", value: "Lovelace", error: "Error: Last Name is required" },
{ key: "Postal code", testId: "postalCode", value: "10115", error: "Error: Postal Code is required" },
] as const;
/** Registers one checkout test for one row of test data. */
export function checkoutTest(title: string, c: Case, tag: string[] = []) {
test(title, { tag }, async ({ page }) => {
await page.goto("/");
await page.getByTestId("username").fill(c.Account);
await page.getByTestId("password").fill(PASSWORD);
await page.getByTestId("login-button").click();
// performance_glitch_user logs in slowly on purpose; the default 5-second expect timeout isn't enough.
await expect(page.getByTestId("inventory-list"), "product list after login").toBeVisible({ timeout: 15_000 });
const item = page.getByTestId("inventory-item").filter({ hasText: c.Product });
const listedPrice = await item.getByTestId("inventory-item-price").textContent();
await item.getByRole("button", { name: "Add to cart" }).click();
await expect(page.getByTestId("shopping-cart-badge"), "cart badge after Add to cart").toHaveText("1");
await page.getByTestId("shopping-cart-link").click();
await page.getByTestId("checkout").click();
for (const field of fields) {
if (c[field.key] === "filled") await page.getByTestId(field.testId).fill(field.value);
}
await page.getByTestId("continue").click();
const firstEmpty = fields.find((field) => c[field.key] === "empty");
if (firstEmpty) {
await expect(page.getByTestId("error"), "validation message").toHaveText(firstEmpty.error);
return;
}
await expect(page.getByTestId("title"), "checkout overview page").toHaveText("Checkout: Overview");
await expect(page.getByTestId("inventory-item-price"), "overview price matches the product list").toHaveText(listedPrice ?? "");
await page.getByTestId("finish").click();
await expect(page.getByTestId("complete-header"), "order confirmation").toHaveText("Thank you for your order!");
});
}The price check compares the price on the product list with the price on the checkout overview. Neither page says what the price should be, but the two must agree.
One test per generated case
import { test, expect } from "@playwright/test";
import cases from "../cases.json";
import { checkoutTest, type Case } from "./checkout";
const PASSWORD = process.env.SAUCE_PASSWORD ?? "secret_sauce";
const viewports = {
desktop: { width: 1280, height: 720 },
mobile: { width: 390, height: 844 },
};
for (const [index, c] of (cases as Case[]).entries()) {
test.describe(`case ${index + 1}`, () => {
// browserName can't be set inside a describe, so each browser project selects its cases by tag.
test.use({ viewport: viewports[c.Viewport as keyof typeof viewports] });
checkoutTest(
`${c.Account} buys ${c.Product} (${c.Viewport}, first ${c["First name"]}, last ${c["Last name"]}, postal ${c["Postal code"]})`,
c,
[`@${c.Browser}`],
);
});
}
test("locked_out_user is refused at login", async ({ page }) => {
await page.goto("/");
await page.getByTestId("username").fill("locked_out_user");
await page.getByTestId("password").fill(PASSWORD);
await page.getByTestId("login-button").click();
await expect(page.getByTestId("error")).toHaveText("Epic sadface: Sorry, this user has been locked out.");
});Choosing the browser per case
Our first version set the browser the same way as the viewport, with test.use({ browserName: c.Browser }) inside each describe. Playwright refused to load the file:
Cannot use({ browserName }) in a describe group, because it forces a new worker.
Make it top-level in the test file or put in the configuration file.The browser belongs to the worker, not to a single test. The fix is one project per browser, each selecting its cases by tag. Untagged tests, such as the locked-out check, run on Chromium:
import { defineConfig } from "@playwright/test";
export default defineConfig({
testDir: "./tests",
fullyParallel: true,
workers: 4,
retries: 0,
reporter: [["list"], ["json", { outputFile: process.env.JSON_OUT ?? "results/results.json" }]],
use: { baseURL: "https://www.saucedemo.com", testIdAttribute: "data-test" },
projects: [
// Untagged tests (the locked-out check and the exhaustive run) go to Chromium.
{ name: "chromium", use: { browserName: "chromium" }, grepInvert: /@firefox|@webkit/ },
{ name: "firefox", use: { browserName: "firefox" }, grep: /@firefox/ },
{ name: "webkit", use: { browserName: "webkit" }, grep: /@webkit/ },
],
});retries: 0 is deliberate. A retry would hide exactly the kind of intermittent failure described later in this article.
Step 4: The results
The 31 cases plus the locked-out test took 71 seconds with four workers on our machine: 19 passed and 13 failed. The 13 failures come from six different bugs:
| Bug | Cases that exposed it | What the test reported |
|---|---|---|
problem_user can't add the Bolt T-Shirt, Fleece Jacket, or red T-Shirt |
10, 12, 20 | Cart badge never appeared |
problem_user: typing a last name overwrites the first name |
23, 24 | Still on "Checkout: Your Information" after Continue |
error_user can't add the same three products |
4, 6, 28 | Cart badge never appeared |
error_user doesn't require a last name |
14 | No validation message; checkout continued |
error_user can't finish an order |
27 | No "Thank you for your order!" after Finish |
visual_user sees different prices on the product list and at checkout |
29, 30, 31 | Overview showed $9.99 and $15.99, not the listed price |
standard_user and performance_glitch_user passed every case, and so did the locked-out test. None of these six bugs depended on the browser or the viewport; the one browser-specific failure we saw was intermittent, and is described below.
Step 5: Check against every combination
Pairwise testing claims that covering pairs is enough. To test that claim, we ran every combination of the parameters that produced failures: 5 accounts × 6 products × 8 form combinations = 240 checkouts. Browser and viewport were left out to keep the run reasonable for a free public site, so this run used Chromium at the default viewport. The exhaustive file reuses the same checkoutTest:
import { checkoutTest } from "./checkout";
// Ground truth: every account × product × form combination, on Chromium at the default viewport.
const accounts = ["standard_user", "problem_user", "performance_glitch_user", "error_user", "visual_user"];
const products = [
"Sauce Labs Backpack",
"Sauce Labs Bike Light",
"Sauce Labs Bolt T-Shirt",
"Sauce Labs Fleece Jacket",
"Sauce Labs Onesie",
"Test.allTheThings() T-Shirt (Red)",
];
const states = ["filled", "empty"] as const;
for (const Account of accounts)
for (const Product of products)
for (const first of states)
for (const last of states)
for (const postal of states)
checkoutTest(`${Account} | ${Product} | ${first} | ${last} | ${postal}`, {
Account,
Product,
"First name": first,
"Last name": last,
"Postal code": postal,
});All 240 ran in 344 seconds with four workers: 165 passed and 75 failed. Grouped by cause, the 75 failures are the same six bugs, and no others:
| Bug | Exhaustive failures | Pairwise cases that exposed it |
|---|---|---|
problem_user can't add three products |
24 | 3 |
problem_user last name overwrites first name |
12 | 2 |
error_user can't add three products |
24 | 3 |
error_user doesn't require a last name |
6 | 1 |
error_user can't finish an order |
3 | 1 |
visual_user price mismatch |
6 | 3 |
The exhaustive run showed some bugs in forms the pairwise run didn't. For problem_user, an empty first name or postal code produced "Error: Last Name is required", because the typed last name had landed in the first name field. For error_user, an empty last name and postal code produced "Error: Postal Code is required", because the last name was never checked. Both are the same bugs seen from another angle.
So on this app, the 31 pairwise cases found every bug that the 240 exhaustive runs found. There's one limit to that claim: the exhaustive run didn't vary the browser or the viewport, so it can't show a browser-specific bug that the pairwise set missed.
What pairwise coverage didn't guarantee
Finding all six bugs is a good result, but two of them were found for a reason pairwise testing doesn't promise.
Two bugs needed four values at once
The error_user Finish bug and the visual_user price bug only appear on the overview page. A case reaches that page only when the first name, last name, and postal code are all filled. That's an interaction of four values: the account and three form fields. Pairwise coverage guarantees every pair, such as visual_user with a filled last name, but not a case where all four line up.
In our set, those cases exist because most rows have no empty field: there are only three negative values to place, and 30 account-product pairs to cover. That's a property of this model, not a promise. A small change to the model can produce a set with no happy path for some account, and the second round later in this article did exactly that.
The fix costs almost nothing. Add the happy path for each account as a required case:
Account = standard_user, First name = filled, Last name = filled, Postal code = filled
Account = problem_user, First name = filled, Last name = filled, Postal code = filled
Account = performance_glitch_user, First name = filled, Last name = filled, Postal code = filled
Account = error_user, First name = filled, Last name = filled, Postal code = filled
Account = visual_user, First name = filled, Last name = filled, Postal code = filledWith these in the generator's "Required cases" box, the generator produced 30 cases: one fewer than before, and the smallest possible for this model. The happy path for every account is now guaranteed, not incidental. Seeding a greedy generator with good rows can shrink the result as well as constrain it.
A failure hides the rest of its case
Case 10 was problem_user buying the Bolt T-Shirt with an empty postal code. It failed at Add to cart, so the empty postal code was never submitted. Each case covers 21 pairs, and a case that stops early only exercises the pairs it reached.
We counted: six cases stopped at Add to cart, and 9 of the 196 pairs appeared only in those cases, so they were never actually tested. Among them were problem_user with an empty first name and problem_user with an empty postal code. Six of the nine pairs combine an account or product with a form value, so the exhaustive run covered them too, and it found only bugs that other cases had already exposed. The other three involve a browser, which the exhaustive run didn't vary. Without an exhaustive run, which you normally won't have, you can't know whether an untested pair hides something.
The fix is to rerun after reporting the bugs. We added one rule to the model, so the known-broken products are no longer combined with the two accounts that can't buy them:
NEVER Account IN problem_user, error_user AND Product IN Sauce Labs Bolt T-Shirt, Sauce Labs Fleece Jacket, Test.allTheThings() T-Shirt (Red)The generator excluded those 6 account-product pairs and produced 26 cases. With the locked-out test, 21 passed and 6 failed:
- Cases 7 and 11 ran
problem_userwith an empty postal code and an empty first name, two of the pairs the first round never tested. Both got "Error: Last Name is required", the name overwrite bug in a new form. - Case 8 showed again that
error_userisn't asked for a last name. - Cases 25 and 26 showed the
visual_userprice mismatch. - Case 2 timed out loading the home page on Firefox. Repeated five times, it passed every time, so we count it as a network failure, not a bug.
Round two didn't expose the error_user Finish bug. None of its four error_user cases had every field filled, so none reached the Finish button. This is the four-value problem from the previous section, happening for real one rule away from a set that caught it. With the five happy-path required cases from above, it couldn't have been missed.
Keep a separate test for each reported bug, so the rule doesn't quietly hide the bug from the suite once it's fixed or reappears.
A flaky failure on one browser
In our second full run, case 5 failed: visual_user on WebKit, with the first name filled and the last name empty, got "Error: First Name is required" instead of the last-name error. The next run passed. We repeated just that case, changing one factor at a time:
| Account and browser | Runs | Failed |
|---|---|---|
visual_user on WebKit |
20 | 2 |
standard_user on WebKit |
20 | 0 |
visual_user on Chromium |
20 | 0 |
visual_user on Firefox |
20 | 0 |
A diagnostic version that read the field right after fill() returned found it empty in 3 of 30 WebKit runs. In the other 27 it held "Ada". We didn't find the cause, so we don't claim it's a product bug or a Playwright bug. We only know it's specific to that account and browser on our machine.
It is also a two-value interaction, account × browser, that pairwise testing did cover, in exactly one case. That case failed in 2 of 20 runs, so a single run will usually miss it. Pairwise testing picks which combinations to run. It can't make an intermittent failure show up.
Our own test failed first
Our first run had three failures that weren't Sauce Demo bugs. One was the network: a Firefox case timed out waiting for the home page to load, and passed when it ran again. The same kind of Firefox timeout happened once more in round two. The other two were our mistakes:
performance_glitch_userfailed on WebKit. Its login took 5.1 seconds when we timed it, just over Playwright's default 5-second assertion timeout. The fix is the explicit 15-second timeout on the first assertion, not a fixed wait.- Two failures reported the wrong step.
problem_usercases failed at "overview price" because the test didn't check which page it was on. Adding the "Checkout: Overview" title assertion made the failure say what actually happened: the form never submitted.
When pairwise testing is the wrong tool
- When values affect each other in sequence. Pairwise testing assumes each case is independent. A bug that needs a specific order of actions, such as adding and removing an item before checkout, isn't a combination of parameter values.
- When one parameter decides everything. If most accounts or roles fail early, the interesting pairs sit behind that failure. Fix or exclude the early failures first, as in round two above.
- When all combinations are cheap. The 240-case run took just under six minutes here. If your full set runs in minutes, run it.
- When the risk is higher than two values. For payments or safety-related settings, use three-way coverage, or required cases for the combinations you know matter.
Conclusion
On the Sauce Demo checkout, 31 pairwise cases on three browsers exposed all six bugs that 240 exhaustive runs found, which is a strong result for a set about 2% the size of all 1,440 combinations. Two of those bugs needed four values at once and were caught because the set happened to include happy paths, so add those as required cases. Cases that fail early leave some pairs untested, so exclude known failures with a rule and generate again. And no combination technique will make a flaky failure appear on demand.
Try it on your own form with our pairwise test case generator, then use our Playwright test reviewer on the data-driven test you build from it.
Sources and further reading
- NIST: Automated Combinatorial Testing for Software
- Kuhn, Wallace, and Gallo: Software Fault Interactions and Implications for Software Testing (IEEE TSE, 2004)
- Microsoft PICT: model syntax, negative values, and seeding
- Playwright: parameterize tests
- Playwright: test options and
test.use - Playwright: projects
- Sauce Demo
- How to write effective test cases
- Flaky tests hide behind retries
- Try it: pairwise test case generator