Introduction
Most guides to Playwright visual testing agree on the same setup: call toHaveScreenshot(), add a tolerance such as maxDiffPixels: 100 or maxDiffPixelRatio: 0.01 to stop flakes, disable animations, and commit the baselines. The code in those guides is usually correct. What they rarely show is what that tolerance lets through.
So we measured it. Every result in this article comes from a small project run on Playwright 1.63.0, Chromium, Windows 11, with a 1280×720 viewport. We changed one thing per run and recorded the exact output.
The short version:
| Change or setting | Result |
|---|---|
Price changed from $19 to $18, default settings |
Failed: 11 pixels different |
Same change, maxDiffPixels: 100 |
Passed |
Same change, maxDiffPixelRatio: 0.01 |
Passed |
Same change, threshold: 0.5 |
Failed: 11 pixels different |
| CSS spinner, default settings | Passed |
Counter updated by setInterval, default settings |
Failed: "Failed to take two consecutive stable screenshots" |
Same counter, clock paused with page.clock |
Passed |
| Timestamp on the page, no mask | Failed: 149 pixels different |
| Same timestamp, masked | Passed |
How toHaveScreenshot compares
Before tuning anything, it helps to know three facts from the API reference:
- It waits until two consecutive screenshots are identical, then compares the last one with the baseline.
- By default, any pixel difference fails the test.
maxDiffPixelsandmaxDiffPixelRatioare unset. threshold(default0.2) is different: it's how different a single pixel's color must be before it counts as changed. It isn't a pixel count.
Some defaults are already the safe choice. animations is "disabled" and caret is "hide", so adding animations: "disabled" to your config changes nothing.
Experiment 1: a price change is 11 pixels
The test page was a pricing card with a heading, a price, a line of text, and a button:
import { test, expect } from "@playwright/test";
const price = process.env.PRICE ?? "$19";
test("pricing card", async ({ page }) => {
await page.setContent(`
<div class="card">
<h2>Pro plan</h2>
<p class="price">${price} / month</p>
<p>Unlimited test runs, 5 parallel workers.</p>
<button>Start trial</button>
</div>`);
await expect(page).toHaveScreenshot("pricing.png");
});We created the baseline with $19, then ran again with PRICE='$18':
Error: expect(page).toHaveScreenshot(expected) failed
11 pixels (ratio 0.01 of all image pixels) are different.Eleven pixels. The digits 9 and 8 share most of their shape, so only the part of the glyph that differs is counted. The diff image highlights a small patch on the lower left of the digit, and nothing else.
We got the same 11-pixel result with an element screenshot of the card (expect(page.locator(".card")).toHaveScreenshot()), so cropping doesn't make a small text change look bigger.
The tolerances that hide it
Then we added the two tolerances that appear in most guides and reran the same $18 change:
| Setting | Allowed on a 1280×720 page | 11-pixel price change |
|---|---|---|
| none (default) | 0 pixels | Failed |
maxDiffPixels: 100 |
100 pixels | Passed |
maxDiffPixelRatio: 0.01 |
9,216 pixels | Passed |
threshold: 0.5 |
0 pixels, only stronger color changes count | Failed |
maxDiffPixelRatio: 0.01 sounds strict because it's "one percent". On a 1280×720 screenshot, it tolerates 9,216 changed pixels. That's room for a wrong price, a missing icon, or a one-word label change.
Raising threshold didn't hide the change, because text changes are high-contrast: a pixel going from white to dark text is far past any threshold. That makes threshold the better tool for anti-aliasing noise, and pixel-count tolerances the risky one.
The error message rounds up
Look again at the failure: 11 pixels (ratio 0.01 of all image pixels). The real ratio is 11 ÷ 921,600, about 0.00001. Playwright rounds the displayed ratio up to two decimals. In Playwright 1.63 the formatting is Math.ceil(count / (width * height) * 100) / 100, so any failure shows at least 0.01.
This matters because a common fix for a failing visual test is to read the ratio and set maxDiffPixelRatio just above it. Doing that here would set a tolerance roughly 800 times larger than the change that caused the failure. Use the pixel count, never the displayed ratio.
Experiment 2: what "disabled animations" doesn't stop
A CSS spinner with an infinite @keyframes animation passed with default settings. Playwright disables CSS animations and transitions before capturing.
A counter updated by JavaScript is different:
test("js ticker", async ({ page }) => {
await page.setContent(`<p id="t">0</p>
<script>let n = 0; setInterval(() => { document.getElementById("t").textContent = ++n; }, 50);</script>`);
await expect(page).toHaveScreenshot("ticker.png");
});It never produced a baseline. It kept taking screenshots until the 5-second expect timeout:
Error: expect(page).toHaveScreenshot(expected) failed
Timeout: 5000ms
Failed to take two consecutive stable screenshots.
- 115 pixels (ratio 0.01 of all image pixels) are different.
- 191 pixels (ratio 0.01 of all image pixels) are different.
- 203 pixels (ratio 0.01 of all image pixels) are different.animations: "disabled" affects CSS animations, CSS transitions, and Web Animations. Timers, requestAnimationFrame loops, carousels driven by JavaScript, and live data all keep changing.
The fix is to control time rather than add a tolerance. Installing Playwright's fake clock and pausing it made the same page stable, and the test passed on the next run:
test("js ticker, clock paused", async ({ page }) => {
await page.clock.install({ time: new Date("2026-01-01T10:00:00Z") });
await page.setContent(`<p id="t">0</p>
<script>let n = 0; setInterval(() => { document.getElementById("t").textContent = ++n; }, 50);</script>`);
await page.clock.pauseAt(new Date("2026-01-01T10:00:01Z"));
await expect(page).toHaveScreenshot("ticker-frozen.png");
});Our first attempt called only page.clock.install(). That still failed the same way: an installed clock keeps running until you pause it.
Experiment 3: mask what you can't control
A "last updated" timestamp rendered with new Date().toISOString() changed on every run and failed with 149 different pixels. Masking the element fixed it:
await expect(page).toHaveScreenshot("clock.png", {
mask: [page.locator("#updated")],
});A mask paints a solid box (pink #FF00FF by default) over the element in both the baseline and the new screenshot. The box itself is still compared. We checked with a masked badge: changing its text from 3 new to 9 new (same width) passed, and changing it to 128 new messages failed, because the pink box got wider. Masking hides what's inside the box, not a change in its size.
Prefer a paused clock or fixed test data for anything you control, and use masks for what you don't: ads, third-party widgets, user avatars.
Baselines: where they live and when they're written
File names include the platform
Our baselines were written next to the spec file, not into a __screenshots__ folder:
tests/visual.spec.ts-snapshots/pricing-chromium-win32.png
tests/visual.spec.ts-snapshots/pricing-card-chromium-win32.pngThe name is the screenshot name, the project, and the operating system. A baseline created on Windows (win32) or macOS (darwin) is never compared on a Linux runner. CI looks for a -linux file, doesn't find one, and fails.
The visual comparisons guide warns that rendering also varies with OS version, hardware, headless mode, and other factors. The practical rule is to generate baselines in the same environment that runs the comparison. For most teams that's the official Docker image, with a tag that matches your installed @playwright/test version exactly. An image a few versions older ships a different browser build and renders differently. We didn't measure cross-OS pixel differences for this article, so we won't quote a number.
A missing baseline fails, but writes one
With no baseline on disk, the first run failed with exit code 1:
Error: A snapshot doesn't exist at tests\visual.spec.ts-snapshots\pricing-chromium-win32.png, writing actual.It also wrote the file. The next run on the same machine passed against a baseline that nobody reviewed. That's the default mode, --update-snapshots=missing.
With --update-snapshots=none, the run failed with the same message and wrote nothing. Use that in CI, so a missing baseline stays a visible failure instead of becoming an unreviewed file.
We also ran the missing-baseline case with --retries=1. Playwright made one attempt, not two, and the test failed. So retries don't turn a missing baseline green, at least in 1.63.
A configuration that doesn't hide regressions
import { defineConfig, devices } from "@playwright/test";
export default defineConfig({
testDir: "./tests",
// Linux baselines only, generated in the Docker image that matches this Playwright version.
snapshotPathTemplate: "{testDir}/__screenshots__/{testFilePath}/{arg}-{projectName}{ext}",
expect: {
toHaveScreenshot: {
// Absorbs anti-aliasing noise in color, not changed content.
threshold: 0.2,
// No maxDiffPixels or maxDiffPixelRatio by default: an 11-pixel price change must fail.
},
},
use: { viewport: { width: 1280, height: 720 } },
projects: [{ name: "chromium", use: { ...devices["Desktop Chrome"] } }],
});Removing {platform} from snapshotPathTemplate is a deliberate choice. It means "there is one set of baselines, and they come from Linux." Only do it if every run that compares or updates screenshots uses the same Docker image. Otherwise, keep the default template.
When a specific screenshot needs tolerance, add it to that one assertion with a number from the pixel count in the failure, and a comment explaining what noise it absorbs. A tolerance in the global config applies to every screenshot, including the ones where 11 pixels is a wrong price.
Updating baselines in pull requests
Run baseline updates in the same image as CI. This command mounts your project into the image and updates only screenshots that changed:
docker run --rm --ipc=host -v "$(pwd)":/work -w /work mcr.microsoft.com/playwright:v1.63.0-noble \
/bin/bash -c "npm ci && npx playwright test --grep @visual --update-snapshots=changed"In CI, compare only, and upload the diffs when a comparison fails:
name: visual
on: pull_request
jobs:
visual:
runs-on: ubuntu-latest
container:
image: mcr.microsoft.com/playwright:v1.63.0-noble
steps:
- uses: actions/checkout@v4
- run: npm ci
# Never write baselines in CI: a missing baseline must stay a failure.
- run: npx playwright test --grep @visual --update-snapshots=none
- if: failure()
uses: actions/upload-artifact@v4
with:
name: visual-diffs
path: test-results/
retention-days: 14Each failure in test-results/ has three images: -expected, -actual, and -diff. Reviewers approve a change by looking at the diff, then the author commits the updated baseline. GitHub shows changed PNG files side by side in the pull request, so the baseline change is reviewed with the code that caused it.
Checklist
- Leave
maxDiffPixelsandmaxDiffPixelRatiounset globally. A price change was 11 pixels. Add tolerance per assertion, with a comment. - Read the pixel count, not the ratio. The ratio is rounded up to at least
0.01. - Use
thresholdfor anti-aliasing noise. It didn't hide a text change even at0.5. - Pause the clock for JavaScript-driven changes.
animations: "disabled"only affects CSS and Web Animations. - Mask only what you don't control. A masked element that changes size still fails.
- Generate and compare baselines in one environment, with a Docker image tag that matches your Playwright version.
- Run CI with
--update-snapshots=none. The default writes missing baselines.
Sources and further reading
- Playwright: visual comparisons
- Playwright:
toHaveScreenshotoptions - Playwright: clock
- Playwright: command line (
--update-snapshots) - Playwright:
snapshotPathTemplate - Playwright: Docker
- pixelmatch, the comparison library behind
threshold - Flaky tests hide behind retries
- Try it: Playwright test reviewer