Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

Playwright Visual Regression Testing: What the Popular Settings Actually Catch

A measured look at Playwright's toHaveScreenshot: how many pixels a real regression changes, why maxDiffPixels: 100 and maxDiffPixelRatio: 0.01 can hide it, what animations: 'disabled' does not freeze, and how to handle baselines in CI.

QA Vibes EditorialPublished 8 minTested with Playwright 1.63.0, Node.js 20.20.2Revision history ↓

Key takeaways

  • A real price change altered only 11 pixels, so maxDiffPixels: 100 or maxDiffPixelRatio: 0.01 let it pass.
  • Leave pixel tolerances unset globally, and use threshold for anti-aliasing noise.
  • Disabled animations don't freeze JavaScript-driven changes; pause the clock with page.clock.
  • Create and compare baselines in the same environment, and run CI with --update-snapshots=none.
Contents (10 sections)

Introduction

Most guides to Playwright visual testing agree on the same setup: call toHaveScreenshot(), add a tolerance such as maxDiffPixels: 100 or maxDiffPixelRatio: 0.01 to stop flakes, disable animations, and commit the baselines. The code in those guides is usually correct. What they rarely show is what that tolerance lets through.

So we measured it. Every result in this article comes from a small project run on Playwright 1.63.0, Chromium, Windows 11, with a 1280×720 viewport. We changed one thing per run and recorded the exact output.

The short version:

Change or setting Result
Price changed from $19 to $18, default settings Failed: 11 pixels different
Same change, maxDiffPixels: 100 Passed
Same change, maxDiffPixelRatio: 0.01 Passed
Same change, threshold: 0.5 Failed: 11 pixels different
CSS spinner, default settings Passed
Counter updated by setInterval, default settings Failed: "Failed to take two consecutive stable screenshots"
Same counter, clock paused with page.clock Passed
Timestamp on the page, no mask Failed: 149 pixels different
Same timestamp, masked Passed

How toHaveScreenshot compares

Before tuning anything, it helps to know three facts from the API reference:

  1. It waits until two consecutive screenshots are identical, then compares the last one with the baseline.
  2. By default, any pixel difference fails the test. maxDiffPixels and maxDiffPixelRatio are unset.
  3. threshold (default 0.2) is different: it's how different a single pixel's color must be before it counts as changed. It isn't a pixel count.

Some defaults are already the safe choice. animations is "disabled" and caret is "hide", so adding animations: "disabled" to your config changes nothing.

Experiment 1: a price change is 11 pixels

The test page was a pricing card with a heading, a price, a line of text, and a button:

import { test, expect } from "@playwright/test";
 
const price = process.env.PRICE ?? "$19";
 
test("pricing card", async ({ page }) => {
  await page.setContent(`
    <div class="card">
      <h2>Pro plan</h2>
      <p class="price">${price} / month</p>
      <p>Unlimited test runs, 5 parallel workers.</p>
      <button>Start trial</button>
    </div>`);
  await expect(page).toHaveScreenshot("pricing.png");
});

We created the baseline with $19, then ran again with PRICE='$18':

Error: expect(page).toHaveScreenshot(expected) failed
  11 pixels (ratio 0.01 of all image pixels) are different.

Eleven pixels. The digits 9 and 8 share most of their shape, so only the part of the glyph that differs is counted. The diff image highlights a small patch on the lower left of the digit, and nothing else.

We got the same 11-pixel result with an element screenshot of the card (expect(page.locator(".card")).toHaveScreenshot()), so cropping doesn't make a small text change look bigger.

The tolerances that hide it

Then we added the two tolerances that appear in most guides and reran the same $18 change:

Setting Allowed on a 1280×720 page 11-pixel price change
none (default) 0 pixels Failed
maxDiffPixels: 100 100 pixels Passed
maxDiffPixelRatio: 0.01 9,216 pixels Passed
threshold: 0.5 0 pixels, only stronger color changes count Failed

maxDiffPixelRatio: 0.01 sounds strict because it's "one percent". On a 1280×720 screenshot, it tolerates 9,216 changed pixels. That's room for a wrong price, a missing icon, or a one-word label change.

Raising threshold didn't hide the change, because text changes are high-contrast: a pixel going from white to dark text is far past any threshold. That makes threshold the better tool for anti-aliasing noise, and pixel-count tolerances the risky one.

The error message rounds up

Look again at the failure: 11 pixels (ratio 0.01 of all image pixels). The real ratio is 11 ÷ 921,600, about 0.00001. Playwright rounds the displayed ratio up to two decimals. In Playwright 1.63 the formatting is Math.ceil(count / (width * height) * 100) / 100, so any failure shows at least 0.01.

This matters because a common fix for a failing visual test is to read the ratio and set maxDiffPixelRatio just above it. Doing that here would set a tolerance roughly 800 times larger than the change that caused the failure. Use the pixel count, never the displayed ratio.

Experiment 2: what "disabled animations" doesn't stop

A CSS spinner with an infinite @keyframes animation passed with default settings. Playwright disables CSS animations and transitions before capturing.

A counter updated by JavaScript is different:

test("js ticker", async ({ page }) => {
  await page.setContent(`<p id="t">0</p>
    <script>let n = 0; setInterval(() => { document.getElementById("t").textContent = ++n; }, 50);</script>`);
  await expect(page).toHaveScreenshot("ticker.png");
});

It never produced a baseline. It kept taking screenshots until the 5-second expect timeout:

Error: expect(page).toHaveScreenshot(expected) failed
Timeout: 5000ms
  Failed to take two consecutive stable screenshots.
  - 115 pixels (ratio 0.01 of all image pixels) are different.
  - 191 pixels (ratio 0.01 of all image pixels) are different.
  - 203 pixels (ratio 0.01 of all image pixels) are different.

animations: "disabled" affects CSS animations, CSS transitions, and Web Animations. Timers, requestAnimationFrame loops, carousels driven by JavaScript, and live data all keep changing.

The fix is to control time rather than add a tolerance. Installing Playwright's fake clock and pausing it made the same page stable, and the test passed on the next run:

test("js ticker, clock paused", async ({ page }) => {
  await page.clock.install({ time: new Date("2026-01-01T10:00:00Z") });
  await page.setContent(`<p id="t">0</p>
    <script>let n = 0; setInterval(() => { document.getElementById("t").textContent = ++n; }, 50);</script>`);
  await page.clock.pauseAt(new Date("2026-01-01T10:00:01Z"));
  await expect(page).toHaveScreenshot("ticker-frozen.png");
});

Our first attempt called only page.clock.install(). That still failed the same way: an installed clock keeps running until you pause it.

Experiment 3: mask what you can't control

A "last updated" timestamp rendered with new Date().toISOString() changed on every run and failed with 149 different pixels. Masking the element fixed it:

await expect(page).toHaveScreenshot("clock.png", {
  mask: [page.locator("#updated")],
});

A mask paints a solid box (pink #FF00FF by default) over the element in both the baseline and the new screenshot. The box itself is still compared. We checked with a masked badge: changing its text from 3 new to 9 new (same width) passed, and changing it to 128 new messages failed, because the pink box got wider. Masking hides what's inside the box, not a change in its size.

Prefer a paused clock or fixed test data for anything you control, and use masks for what you don't: ads, third-party widgets, user avatars.

Baselines: where they live and when they're written

File names include the platform

Our baselines were written next to the spec file, not into a __screenshots__ folder:

tests/visual.spec.ts-snapshots/pricing-chromium-win32.png
tests/visual.spec.ts-snapshots/pricing-card-chromium-win32.png

The name is the screenshot name, the project, and the operating system. A baseline created on Windows (win32) or macOS (darwin) is never compared on a Linux runner. CI looks for a -linux file, doesn't find one, and fails.

The visual comparisons guide warns that rendering also varies with OS version, hardware, headless mode, and other factors. The practical rule is to generate baselines in the same environment that runs the comparison. For most teams that's the official Docker image, with a tag that matches your installed @playwright/test version exactly. An image a few versions older ships a different browser build and renders differently. We didn't measure cross-OS pixel differences for this article, so we won't quote a number.

A missing baseline fails, but writes one

With no baseline on disk, the first run failed with exit code 1:

Error: A snapshot doesn't exist at tests\visual.spec.ts-snapshots\pricing-chromium-win32.png, writing actual.

It also wrote the file. The next run on the same machine passed against a baseline that nobody reviewed. That's the default mode, --update-snapshots=missing.

With --update-snapshots=none, the run failed with the same message and wrote nothing. Use that in CI, so a missing baseline stays a visible failure instead of becoming an unreviewed file.

We also ran the missing-baseline case with --retries=1. Playwright made one attempt, not two, and the test failed. So retries don't turn a missing baseline green, at least in 1.63.

A configuration that doesn't hide regressions

import { defineConfig, devices } from "@playwright/test";
 
export default defineConfig({
  testDir: "./tests",
  // Linux baselines only, generated in the Docker image that matches this Playwright version.
  snapshotPathTemplate: "{testDir}/__screenshots__/{testFilePath}/{arg}-{projectName}{ext}",
  expect: {
    toHaveScreenshot: {
      // Absorbs anti-aliasing noise in color, not changed content.
      threshold: 0.2,
      // No maxDiffPixels or maxDiffPixelRatio by default: an 11-pixel price change must fail.
    },
  },
  use: { viewport: { width: 1280, height: 720 } },
  projects: [{ name: "chromium", use: { ...devices["Desktop Chrome"] } }],
});

Removing {platform} from snapshotPathTemplate is a deliberate choice. It means "there is one set of baselines, and they come from Linux." Only do it if every run that compares or updates screenshots uses the same Docker image. Otherwise, keep the default template.

When a specific screenshot needs tolerance, add it to that one assertion with a number from the pixel count in the failure, and a comment explaining what noise it absorbs. A tolerance in the global config applies to every screenshot, including the ones where 11 pixels is a wrong price.

Updating baselines in pull requests

Run baseline updates in the same image as CI. This command mounts your project into the image and updates only screenshots that changed:

docker run --rm --ipc=host -v "$(pwd)":/work -w /work mcr.microsoft.com/playwright:v1.63.0-noble \
  /bin/bash -c "npm ci && npx playwright test --grep @visual --update-snapshots=changed"

In CI, compare only, and upload the diffs when a comparison fails:

name: visual
 
on: pull_request
 
jobs:
  visual:
    runs-on: ubuntu-latest
    container:
      image: mcr.microsoft.com/playwright:v1.63.0-noble
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      # Never write baselines in CI: a missing baseline must stay a failure.
      - run: npx playwright test --grep @visual --update-snapshots=none
      - if: failure()
        uses: actions/upload-artifact@v4
        with:
          name: visual-diffs
          path: test-results/
          retention-days: 14

Each failure in test-results/ has three images: -expected, -actual, and -diff. Reviewers approve a change by looking at the diff, then the author commits the updated baseline. GitHub shows changed PNG files side by side in the pull request, so the baseline change is reviewed with the code that caused it.

Checklist

  • Leave maxDiffPixels and maxDiffPixelRatio unset globally. A price change was 11 pixels. Add tolerance per assertion, with a comment.
  • Read the pixel count, not the ratio. The ratio is rounded up to at least 0.01.
  • Use threshold for anti-aliasing noise. It didn't hide a text change even at 0.5.
  • Pause the clock for JavaScript-driven changes. animations: "disabled" only affects CSS and Web Animations.
  • Mask only what you don't control. A masked element that changes size still fails.
  • Generate and compare baselines in one environment, with a Docker image tag that matches your Playwright version.
  • Run CI with --update-snapshots=none. The default writes missing baselines.

Sources and further reading

Tools mentioned

PlaywrightUI AutomationOpen source
GitHub ActionsCI/CDFree plan

Links go to each tool’s official site. How we choose and link tools

Revision history

First published.

Spotted a mistake? Report it — corrections land here.

Written and reviewed by

QA Vibes Editorial

Articles are written and reviewed by practicing QA and automation engineers. Every article lists its sources and shows when it was last updated.

Check your own tests

Playwright test reviewer

Paste a test to find fixed waits, missing awaits, and assertions that can never fail.