Skip to content
Independent guides for QA & test automationRSSEditorial policy
QA Vibes

How to Run Automated Tests in CI/CD with GitHub Actions: From One Job to Shards

A practical path to tests in CI: the GitHub Actions workflow Playwright generates and what each line does, the additions a real project needs (taken from this site's pipeline), and how to shard a slow suite and merge the results into one report.

QA Vibes EditorialPublished Updated 9 minTested with Playwright 1.63.0Revision history ↓

Key takeaways

  • Start from the workflow Playwright generates with --gha; it is a correct first pipeline.
  • Install browsers with --with-deps, run cheap checks first, and upload reports when tests fail.
  • Give the workflow token only the permissions it needs, and keep credentials in GitHub secrets.
  • Shard a slow suite with the blob reporter, then merge the shards into one report.
  • Make the test check required before merging, so green actually means something.
Contents (11 sections)

Introduction

Tests protect a product only if they run on every change, without anyone remembering to start them. Continuous integration (CI) does that: when someone opens a pull request, a server checks out the code, runs the tests, and marks the change red or green.

This guide builds a GitHub Actions pipeline in three steps:

  1. The workflow Playwright generates, explained line by line.
  2. The additions a real project needs, taken from the CI pipeline that tests this website.
  3. Sharding, which splits a slow suite across machines, plus merging the pieces into one report.

The generated workflow and the sharding commands come from Playwright 1.63.0; we ran the sharding and report merge ourselves. The pipeline in step 2 is the one that runs on every change to this site.

What CI should do with your tests

Before writing YAML, decide what runs when:

When What runs Goal
Every pull request Lint, type checks, unit tests, a build, and the main end-to-end tests Block broken changes before they merge
Every merge to main The same, plus anything too slow for pull requests Confirm main is always releasable
On a schedule (for example, nightly) Long suites, cross-browser runs, checks against third-party services Catch slow drift without slowing down developers

A pull request check should take minutes, not an hour. If it's slow, developers stop waiting for it, and it stops protecting anything.

Step 1: the workflow Playwright generates

Run npm init playwright@latest -- --gha, or answer "yes" when the installer asks about GitHub Actions. With Playwright 1.63.0 it wrote .github/workflows/playwright.yml:

name: Playwright Tests
on:
  push:
    branches: [ main, master ]
  pull_request:
    branches: [ main, master ]
jobs:
  test:
    timeout-minutes: 60
    runs-on: ubuntu-latest
    steps:
    - uses: actions/checkout@v4
    - uses: actions/setup-node@v4
      with:
        node-version: lts/*
    - name: Install dependencies
      run: npm ci
    - name: Install Playwright Browsers
      run: npx playwright install --with-deps
    - name: Run Playwright tests
      run: npx playwright test
    - uses: actions/upload-artifact@v4
      if: ${{ !cancelled() }}
      with:
        name: playwright-report
        path: playwright-report/
        retention-days: 30

What each part does:

  • on runs the workflow on pushes to main and on pull requests that target it.
  • timeout-minutes: 60 stops a stuck job instead of letting it run for GitHub's default of six hours.
  • npm ci installs exactly what's in package-lock.json, unlike npm install, which can update it.
  • npx playwright install --with-deps downloads the browsers and the Linux libraries they need. Forgetting --with-deps is the most common reason Playwright fails on a fresh CI machine but works locally.
  • if: ${{ !cancelled() }} uploads the HTML report even when tests fail, which is exactly when you need it. It skips the upload only if someone cancelled the run.

The generated playwright.config.ts also changes behavior on CI: forbidOnly fails the run if someone committed test.only, and retries is set to 2 when the CI environment variable is present, which GitHub Actions sets automatically.

This workflow is a good start. Commit it, open a pull request, and you have working CI.

Step 2: what a real project adds

This site is a Next.js app tested with Playwright. Here are the parts of its pipeline that go beyond the generated file, each with the problem it solves:

name: CI
 
on:
  push:
    branches: [main]
  pull_request:
 
concurrency:
  group: ci-${{ github.ref }}
  cancel-in-progress: true
 
jobs:
  test:
    runs-on: ubuntu-latest
    timeout-minutes: 20
    permissions:
      contents: read
    steps:
      - uses: actions/checkout@v4
 
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
 
      - run: npm ci
 
      - name: Lint
        run: npm run lint
 
      - name: Typecheck
        run: npm run typecheck
 
      - name: No fixme tests
        run: |
          if grep -rnE "^\s*test(\.describe)?\.fixme\(" tests/; then
            echo "::error::test.fixme() found. Fix the test or file a bug instead of skipping it."
            exit 1
          fi
 
      - name: Build
        run: npm run build
 
      - name: Install Playwright browsers
        run: npx playwright install --with-deps chromium
 
      - name: Unit and end-to-end tests
        run: npm test
 
      - name: Upload Playwright report
        if: failure()
        uses: actions/upload-artifact@v4
        with:
          name: playwright-report
          path: |
            playwright-report/
            test-results/
          retention-days: 14

What each addition is for:

  • concurrency with cancel-in-progress. Pushing three commits in a row to a pull request starts three runs; only the newest matters. This cancels the older ones and saves CI minutes.
  • permissions: contents: read. The workflow's token can read the code and nothing else. If a dependency is ever compromised, it can't push commits or change releases with that token. Grant more only to jobs that need it.
  • cache: npm. Reuses downloaded packages between runs, so npm ci is faster.
  • Pinned Node.js version. lts/* changes when a new LTS release comes out; a fixed version means the pipeline changes only when you change it.
  • Cheap checks first. Lint and type checks fail in seconds, so a typo doesn't wait behind a full build and browser install.
  • A guard against skipped tests. test.fixme silently turns a test off. This step fails the build when one appears, so skipping a test is a visible decision, not an accident.
  • Build before end-to-end tests. The tests run against the production build, the same artifact that gets deployed, not a development server.
  • Only one browser. install --with-deps chromium downloads just the browser the tests use, which is noticeably faster than installing all three.
  • Upload on failure(), with traces. test-results/ contains the trace files, so a failure can be debugged from the artifact without re-running anything.

The real file also runs a report-triage step and keeps a longer list of permissions for it; those parts are left out here.

Secrets

Tests often need credentials: a test account password, an API key, or a cloud grid token.

  1. Add them under Settings → Secrets and variables → Actions in the repository.
  2. Pass them to the step that needs them as environment variables:
      - name: Run tests against staging
        env:
          BASE_URL: https://staging.example.com
          TEST_PASSWORD: ${{ secrets.TEST_PASSWORD }}
        run: npx playwright test
  1. Read them in tests with process.env.TEST_PASSWORD. Never commit them to the repository.

GitHub masks secret values in logs, but only exact matches. Don't print them, and don't write them into files that get uploaded as artifacts, such as traces or screenshots of a login form.

Secrets aren't available to workflows triggered by pull requests from forks, which is intentional: a stranger's pull request shouldn't be able to read your secrets.

Step 3: shard a slow suite

When a suite takes too long on one machine, Playwright can split it into shards: --shard=1/4 runs the first quarter of the tests, --shard=2/4 the second, and so on. Each shard runs on its own machine at the same time.

The catch is reporting: four machines produce four partial reports. Playwright solves this with the blob reporter, which writes a raw report per shard, and merge-reports, which combines them.

Try it locally first

We split 10 tests into two shards on one machine to see how it works:

npx playwright test --shard=1/2 --reporter=blob
npx playwright test --shard=2/2 --reporter=blob

Each run replaces the blob-report folder, so we copied each shard's zip file into a shared all-blob folder after it finished. The first shard ran 5 tests, all passing. The second ran the other 5: 4 passed and 1 failed, a price check on a deliberately broken demo account. Then:

npx playwright merge-reports --reporter=list all-blob

The merged output listed all 10 tests, including the failure's full error and a trace path, and ended with:

  1 failed
    [chromium] › tests\shop\checkout.spec.ts:61:7 › visual_user › prices match the catalog
  9 passed (20.3s)

One report, as if the whole suite had run on one machine.

The same thing on GitHub Actions

This follows the pattern in Playwright's sharding documentation. Set the reporter to blob on CI in playwright.config.ts (reporter: process.env.CI ? "blob" : "html"), then:

jobs:
  test:
    timeout-minutes: 30
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        shardIndex: [1, 2, 3, 4]
        shardTotal: [4]
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps
      - run: npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}
      - uses: actions/upload-artifact@v4
        if: ${{ !cancelled() }}
        with:
          name: blob-report-${{ matrix.shardIndex }}
          path: blob-report
          retention-days: 1
 
  merge-reports:
    if: ${{ !cancelled() }}
    needs: [test]
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - uses: actions/download-artifact@v4
        with:
          path: all-blob-reports
          pattern: blob-report-*
          merge-multiple: true
      - run: npx playwright merge-reports --reporter html ./all-blob-reports
      - uses: actions/upload-artifact@v4
        with:
          name: html-report--attempt-${{ github.run_attempt }}
          path: playwright-report
          retention-days: 14

Two details matter:

  • fail-fast: false lets the other shards finish when one fails. Otherwise GitHub cancels them, and your merged report is missing three quarters of the results.
  • The merge job runs if: ${{ !cancelled() }}, so it builds the report even when tests failed, which is when you need the report most.

Shard only when you need to. For a suite that finishes in a few minutes on one machine, the setup time of extra machines can cost more than it saves. Try more workers on one machine first.

Make the check required

A CI check that can be ignored will be ignored. In Settings → Branches (or Rules → Rulesets), add a rule for main that requires the test job to pass before merging. From then on, a red build blocks the merge button.

When CI is red but the change looks fine

  • It fails only on CI. Download the report artifact and open the trace. Common causes: a missing --with-deps, a different screen size, time zone or locale differences, and missing secrets.
  • It fails, then passes on retry. That's a flaky test, and retries are hiding it. Read flaky tests hide behind retries before you raise the retry count.
  • It's slow. Check where the time goes: dependency install (add caching), browser install (install one browser), or the tests themselves (add workers, then shards).

Checklist

  • Tests run on every pull request, and the check is required before merging.
  • Cheap checks (lint, types, unit tests) run before slow ones.
  • Dependencies install with npm ci and are cached.
  • Browsers install with --with-deps, and only the ones you test.
  • Reports and traces upload when tests fail.
  • The workflow token has only the permissions it needs.
  • Secrets come from GitHub secrets and are never printed.
  • Jobs have a timeout, and outdated runs are cancelled.

Conclusion

Start with the workflow Playwright generates; it's correct and it's enough for a first pipeline. Then add what your project needs: fast checks first, least-privilege permissions, a guard against silently skipped tests, and reports you can debug from. When the suite outgrows one machine, shard it and merge the blob reports into one. Finally, make the check required, so green actually means something.

Sources and further reading

Tools mentioned

GitHub ActionsCI/CDFree plan
PlaywrightUI AutomationOpen source
BrowserStackDevice CloudFree trial
SlackReporting & MonitoringFree plan

Links go to each tool’s official site. How we choose and link tools

Revision history

Updated source links that had moved: the pages still exist, at new addresses.
Rewritten around the workflow Playwright 1.63.0 generates, this site's real CI pipeline, and a sharding and report merge we ran.
Revised during a site-wide content audit.
First published.

Spotted a mistake? Report it — corrections land here.

Written and reviewed by

QA Vibes Editorial

Articles are written and reviewed by practicing QA and automation engineers. Every article lists its sources and shows when it was last updated.

Check your own tests

Playwright test reviewer

Paste a test to find fixed waits, missing awaits, and assertions that can never fail.