Introduction
AI-assisted test generation is the most discussed testing topic of 2026, and most articles about it describe what agents could do. This one describes what Playwright's own agents are, based on running the setup command and reading every file it produced.
Playwright ships three test agents: a planner that explores an app and writes a Markdown test plan, a generator that turns that plan into Playwright tests, and a healer that runs failing tests and repairs them. They are not a separate product. They are agent definitions plus an MCP server that ship inside the @playwright/test package you already install.
Everything below was produced with Playwright 1.63.0 in an empty project, using the Claude Code loop. File contents are quoted from that run.
Prerequisites
- An existing Playwright Test project (
playwright.config.tsand at least one test). - An agent client. The CLI in 1.63 accepts
claude,codex,copilot,opencode,vscode, andvscode-legacyas loops. - For VS Code, the Playwright docs state that VS Code 1.105 or newer is required.
Check what your installed version supports before following any tutorial, including this one:
npx playwright --version
npx playwright init-agents --helpRun the setup
Our scratch project had one config file and one seed test:
// tests/seed.spec.ts
import { test } from "@playwright/test";
test("seed", async ({ page }) => {
await page.goto("https://playwright.dev");
});Then:
npx playwright init-agents --loop=claudeThe command printed the files it created:
📝 specs\README.md - directory for test plans
🤖 .claude\agents\playwright-test-generator.md - agent definition
🤖 .claude\agents\playwright-test-healer.md - agent definition
🤖 .claude\agents\playwright-test-planner.md - agent definition
🔧 .mcp.json - mcp configuration
✅ Done.Nothing else in the project changed. The seed test and config were left untouched.
File by file
.mcp.json: the only moving part
{
"mcpServers": {
"playwright-test": {
"command": "cmd",
"args": ["/c", "npx", "playwright", "run-test-mcp-server"]
}
}
}All three agents talk to the browser and the test runner through this one MCP server, started from your local Playwright install. We ran the command on Windows, which is why it wraps npx in cmd /c. Because the server comes from your installed package, the agents always use the same Playwright version as your test suite. That is also why the docs tell you to re-run init-agents after upgrading Playwright: the agent instructions and tool list change between releases.
The seed test
The seed test is how you tell agents where to start. Agents run it to get a page into a known state before they explore or generate anything. If your app needs a login, fixtures, or seeded data, put that setup in the seed test. A good seed test does not assert anything interesting. It only produces a ready page.
playwright-test-planner.md
The planner's front matter gives it read-only project access (Glob, Grep, Read, LS) and browser tools such as browser_navigate, browser_click, browser_type, and browser_snapshot, plus two planner-specific tools: planner_setup_page and planner_save_plan.
Its instructions tell it to set up the page once, explore using accessibility snapshots rather than screenshots ("Do not take screenshots unless absolutely necessary"), map user flows, and save a plan. Plans land in specs/ as Markdown with steps and expected results. The planner cannot edit your code. Its output is a document for a human to review.
playwright-test-generator.md
The generator gets the same read-only project access and a narrower set of browser tools, including verification tools like browser_verify_element_visible and browser_verify_text_visible. It also gets generator_setup_page, generator_read_log, and generator_write_test.
Its workflow is specific. For each scenario in a plan it performs every step live in a real browser, reads the log of what actually worked, then writes one test per file. The instructions require the describe block to match the plan item and a comment with the step text before each step. Tests are recorded from real interactions rather than guessed from source code, which is the main reason to prefer this over asking a general chatbot to "write Playwright tests".
playwright-test-healer.md
The healer has the most power of the three. Beyond read access, it has Edit, MultiEdit, and Write, plus test-runner tools: test_list, test_run, and test_debug. Its workflow: run all tests, debug each failure, inspect the page and console, find the root cause, edit the test, and re-run until it passes.
Three details in the generated instructions matter more than the rest:
- Its list of possible root causes includes "application changes that broke test assumptions", but its remediation step only edits test code, and explicitly includes fixing assertions and expected values.
- It is told it is not interactive: it must not ask the user questions and should do the most reasonable thing to pass the test.
- If a failure persists and it is confident the test itself is correct, it marks the test
test.fixme(), which skips it, and adds a comment explaining what happened.
The healer needs guardrails
A failing test means either the test is broken or the app is. The healer can only change tests. So a real regression has two ways out of a healing run, and neither one is a red build:
- The assertion gets updated to match the new, wrong behavior. This happens when the regression looks like a changed string or value.
- The test gets marked
fixme. This happens when the healer concludes the app is at fault. The regression is now a skipped test with a comment in it.
Both outcomes are reasonable when a person reviews them. Both are dangerous when nobody does. Before letting a healer run on a suite you depend on, we recommend:
- Run it only on a branch and merge through normal code review. Never let it push to your main branch.
- Treat new
test.fixme()calls as bug reports. Fail CI or require sign-off when the number offixmeorskipmarkers goes up. - Review assertion changes separately from locator changes. A changed locator is usually maintenance. A changed or deleted
expectis a behavior decision that needs a human. - Link every heal to a reason. If the app changed on purpose, the pull request should point to that change. If nobody can name one, treat the failure as a bug, not a test fix.
- Scrutinize timing fixes. An added wait from an agent deserves the same review you would give a human adding
waitForTimeout.
To apply these checks to a real change, paste the original and healed versions of a test into our Playwright test reviewer. Its compare mode flags removed assertions, changed expected text, new fixme or skip calls, and added fixed waits.
The planner and generator carry less risk because their output is new files. Review generated tests the way you would review a new teammate's first pull request: check that each test asserts something meaningful, not just that the page loaded.
Where agents fit in a normal workflow
- Planner: useful for exploring an unfamiliar area of an app and producing a checklist you can edit. Treat its plan as a draft for test design, not a coverage guarantee.
- Generator: useful for turning a reviewed plan into first-draft tests quickly, especially for straightforward flows.
- Healer: useful after deliberate UI changes that break many locators at once, as long as the guardrails above apply.
- Not a replacement for: deciding what is worth testing, choosing assertions that reflect business rules, or the test pyramid. Agents make end-to-end tests cheaper to write, which makes it easier to write too many of them.
Conclusion
init-agents is a small, inspectable setup: three Markdown agent definitions, one MCP server config, and a folder for plans. Reading those files tells you exactly what each agent can touch. The planner only reads and documents, the generator records tests from real browser actions, and the healer can edit your tests, which is where human review matters most.
Try it in a throwaway branch, read the generated definitions before trusting them, and re-run init-agents whenever you upgrade Playwright.