Introduction
When an automated test fails for no obvious reason, people blame the locator or the network. Often the cause is data: another test changed the record this one expected, a previous run left a duplicate behind, or someone on the team deleted the "test user" everyone shares.
This guide shows the problem on a real shared system, then fixes it with code you can copy. The examples use Restful Booker, a public practice API for hotel bookings, and Playwright's API testing support. The ideas apply to any framework.
The problem: shared data is somebody else's data
Restful Booker is shared by everyone who practices on it, which makes it a good model of a team's staging environment. When we searched it for bookings with the first name "Test", it returned 4 bookings we had not created.
Imagine a test that creates a booking for "Test" and then asserts the search returns one result. It passes on an empty database, fails on staging, and passes again after someone cleans up. Nothing about the product changed. That's a flaky test caused entirely by data.
Three rules prevent it:
- Own your data. Each test creates what it needs instead of relying on records that already exist.
- Make it unique. Two tests running at the same time, or the same test run twice, never produce the same identifying values.
- Clean it up. Delete what the test created, including when the test fails.
Rules 1 and 2: a builder with unique defaults
A builder returns a valid object with sensible defaults. Each test overrides only the fields it cares about, which keeps tests short and makes their intent obvious.
import { test as base, expect, type APIRequestContext } from "@playwright/test";
export type Booking = {
firstname: string;
lastname: string;
totalprice: number;
depositpaid: boolean;
bookingdates: { checkin: string; checkout: string };
additionalneeds?: string;
};
/** Valid defaults; each test overrides only the fields it cares about. */
export function buildBooking(overrides: Partial<Booking> = {}): Booking {
return {
firstname: "Test",
lastname: `run-${test.info().workerIndex}-${Date.now()}`,
totalprice: 100,
depositpaid: true,
bookingdates: { checkin: "2026-10-01", checkout: "2026-10-03" },
...overrides,
};
}The last name is what makes each booking unique. It combines the Playwright worker index, which differs between tests running in parallel, with a timestamp, which differs between runs. Use whatever field your system searches or deduplicates on: an email address, an order reference, or a username.
This is the first half of tests/data/fixtures.ts. The second half is the fixture.
Rule 3: cleanup that runs even when the test fails
Cleanup in afterEach has two problems: it has to know what the test created, and it's easy to forget. A fixture solves both. It hands the test a function that creates bookings, remembers every ID, and deletes them after the test finishes, pass or fail.
async function token(request: APIRequestContext): Promise<string> {
const response = await request.post("/auth", {
data: { username: "admin", password: process.env.BOOKER_PASSWORD ?? "password123" },
});
return (await response.json()).token;
}
type Fixtures = {
createBooking: (overrides?: Partial<Booking>) => Promise<{ id: number; booking: Booking }>;
};
export const test = base.extend<Fixtures>({
baseURL: "https://restful-booker.herokuapp.com",
createBooking: async ({ request }, use) => {
const created: number[] = [];
await use(async (overrides) => {
const booking = buildBooking(overrides);
const response = await request.post("/booking", { data: booking });
expect(response.ok()).toBeTruthy();
const { bookingid } = await response.json();
created.push(bookingid);
return { id: bookingid, booking };
});
// Teardown runs after the test, even when the test failed.
const auth = await token(request);
for (const id of created) {
await request.delete(`/booking/${id}`, { headers: { Cookie: `token=${auth}` } });
}
},
});
export { expect };Everything after await use(...) is teardown. Playwright runs it when the test ends, whether the test passed, failed an assertion, or threw an error. Because the fixture records IDs as it creates them, cleanup deletes exactly what this test made and nothing else.
Tests that use the fixture
tests/data/bookings.spec.ts:
import { test, expect } from "./fixtures";
test("booking without a deposit keeps depositpaid false", async ({ createBooking, request }) => {
const { id } = await createBooking({ depositpaid: false });
const response = await request.get(`/booking/${id}`);
expect((await response.json()).depositpaid).toBe(false);
});
test("booking with additional needs stores them", async ({ createBooking, request }) => {
const { id } = await createBooking({ additionalneeds: "Late checkout" });
const response = await request.get(`/booking/${id}`);
expect((await response.json()).additionalneeds).toBe("Late checkout");
});
test("search by name finds only this run's booking", async ({ createBooking, request }) => {
const { id, booking } = await createBooking();
const response = await request.get("/booking", {
params: { firstname: booking.firstname, lastname: booking.lastname },
});
const ids = (await response.json()).map((b: { bookingid: number }) => b.bookingid);
expect(ids).toEqual([id]);
});
test("the shared name 'Test' already matches other people's data", async ({ request }) => {
const response = await request.get("/booking", { params: { firstname: "Test" } });
const matches = await response.json();
test.info().annotations.push({ type: "matches", description: String(matches.length) });
expect(Array.isArray(matches)).toBe(true);
});Each test states only what's special about its data: no deposit, or a note. The third test asserts that the search returns exactly its own booking, which is only safe because the last name is unique. Searching by first name alone would have included those 4 other bookings. The last test records how many shared "Test" bookings exist, as an annotation in the report.
Our run, with four tests in parallel against the shared API:
Running 4 tests using 4 workers
ok 3 [chromium] › tests\data\bookings.spec.ts:27:5 › the shared name 'Test' already matches other people's data (689ms)
ok 2 [chromium] › tests\data\bookings.spec.ts:10:5 › booking with additional needs stores them (1.1s)
ok 1 [chromium] › tests\data\bookings.spec.ts:3:5 › booking without a deposit keeps depositpaid false (1.1s)
ok 4 [chromium] › tests\data\bookings.spec.ts:17:5 › search by name finds only this run's booking (1.1s)
4 passed (2.6s)When you can't create data through an API
Creating data through the product's own API is the best default: it's fast, it goes through the same validation as real users, and cleanup is an API call. It isn't always possible. Here are the usual alternatives:
| Strategy | Good for | Watch out for |
|---|---|---|
| Create through the API per test | Most transactional data: users, orders, bookings | Needs an API for every entity, including delete |
| Seed the database directly | Legacy systems without APIs; bulk setup | Bypasses validation, so data can be impossible in real life; couples tests to the schema |
| Restore a snapshot before a run | Large, realistic datasets; reporting and search tests | Slow to refresh; tests in the same run still share data |
| Throwaway database per run (containers) | Integration tests in CI | Needs infrastructure; doesn't help with shared staging |
| Static reference data in version control | Countries, currencies, product catalogs | Only for data that tests read and never change |
Mixing is normal: reference data from version control, a disposable database in CI, and per-test creation for everything a test modifies.
Don't copy production data into test environments
Production data is realistic, which is why teams are tempted to copy it. It also contains real names, emails, addresses, and payment details, and test environments are rarely protected like production.
Safer options, in order of preference:
- Synthetic data, generated to look real without describing real people. Libraries like Faker produce plausible names and addresses.
- De-identified data, where direct identifiers are removed or replaced. NIST SP 800-188 describes the techniques and their limits. De-identification is harder than it looks: combinations of fields such as postcode, birth date, and gender can re-identify people.
- Masking and subsetting tools (for example Tonic or Delphix), when you need production-like volume and relationships and have the budget and a data protection review.
Whatever you choose, never put real credentials in test data files. Load them from environment variables, as the fixture above does with BOOKER_PASSWORD.
Checklist
- No test depends on a record it didn't create, except read-only reference data.
- Identifying fields (emails, names, references) are unique per test and per run.
- Cleanup lives in a fixture or equivalent that runs even when the test fails.
- Tests assert on their own data, never on "the first result" of a shared list.
- The suite passes when run twice in a row, and when run in parallel.
- No production personal data is copied into test environments without de-identification and review.
- Secrets come from environment variables or a secret store.
The fifth point is the cheapest check of all: run the suite twice. If the second run fails, the first one left something behind.
Conclusion
Most "flaky" data problems disappear when every test owns its data, makes it unique, and cleans it up. A builder keeps the defaults in one place; a fixture guarantees the cleanup. On shared environments like staging, and like the public API used here, that's the difference between tests that pass on your machine and tests that pass everywhere.