Skip to content

Playwright Test Agents: How Planner, Generator, Healer Work

ILIA KARPENKO5 min read

Playwright Test Agents are three AI agent definitions that ship with Playwright since version 1.56: a planner that explores your app and writes a Markdown test plan, a generator that turns that plan into Playwright test files, and a healer that runs failing tests and patches them. You can use each one on its own or chain all three, and they run inside the coding agent you already use, such as Claude Code, VS Code with Copilot, Codex, or OpenCode.

Most teams that try AI for end-to-end tests hit the same wall: one prompt asks the model to understand the app, decide what to test, write the code, and debug it, all at once. The agents split that into three jobs with a file between each, so a person can read the plan before any code exists and read the diff before a fix lands.

How to set up Playwright Test Agents

You need a project with @playwright/test installed. From the project root, generate the agent definitions for your tool:

ToolCommand
Claude Codenpx playwright init-agents --loop=claude
VS Code (1.105 or newer)npx playwright init-agents --loop=vscode
Codexnpx playwright init-agents --loop=codex
OpenCodenpx playwright init-agents --loop=opencode

The command writes the agent instructions and the MCP tool configuration they rely on. Under the hood, each agent is a set of instructions plus Playwright MCP tools. Rerun the command whenever you upgrade Playwright, because the definitions change with the tools.

Then add a seed test. It is an ordinary spec that sets up whatever the app needs before a scenario starts: global setup, fixtures, a logged-in user.

import { test, expect } from './fixtures';

test('seed', async ({ page }) => {
  // uses your custom fixtures; agents copy this setup into every test
});

The agents run the seed first and reuse it as the template for every test they generate, so a seed that logs in through the UI on every run makes every generated test do the same.

Free and open sourceCasely writes these cases for youAttach a spec and one file of your team's existing test cases in Claude. Casely copies your columns, names the gaps it found in the spec, and exports a single Excel file your tracker imports in one pass.Read the install docs

What each agent takes in and hands back

AgentInputOutput
PlannerA request ("plan guest checkout"), the seed test, optionally a PRDA Markdown plan in specs/ with scenarios, steps, and expected results
GeneratorA plan from specs/Test files in tests/, one per scenario where possible
HealerThe name of a failing testA passing test, or the test marked skipped if the feature looks broken

The planner opens the app through the seed, walks the flow you asked about, and writes scenarios like "Add valid todo: click the input, type Buy groceries, press Enter; expect the todo in the list, the counter to read 1 item left, and the input to be empty." The plan is plain Markdown, so you can edit it like any document.

The generator works through the plan in a live browser and checks each locator and assertion as it goes, then writes the spec. Each generated file carries comments pointing back to its plan and seed, and each step from the plan appears as a comment above the code that performs it.

The healer replays a failing test, looks at the current page for the element or flow that moved, proposes a patch such as a new locator or a different wait, and reruns until the test passes or its guardrails stop it.

The files end up in a predictable layout:

repo/
  specs/                 # Markdown test plans
  tests/                 # generated Playwright tests
    seed.spec.ts         # seed test
  playwright.config.ts

Where the loop breaks

The planner tests what it can see. It discovers scenarios by exploring the running app. Rules that never show up on screen, such as a rate limit, a timeout, or a permission enforced only by the API, are missing from the plan unless you hand it a PRD or requirements that state them. Attaching the PRD is optional in the docs and close to mandatory in practice.

Plans describe current behavior. An agent exploring a buggy build writes expected results that match the bug. "Counter shows 2 items left" after adding one todo is a valid-looking line in a Markdown plan. A person who knows the requirement has to catch it at the plan stage, because the generator will faithfully assert it.

The healer can hide a regression. Its job is to make a failing test pass. When a locator moved because a button was renamed, that is the right fix. When the test failed because the feature broke, a patch that changes the assertion turns a real failure green. The docs say the healer skips a test it believes is broken; review every healer diff anyway, and treat a changed expected value as a question, not a fix. A test the healer skips as broken deserves a proper bug report, and one that fails only sometimes is a flaky test, which no locator patch will cure.

Generated tests inherit the seed. Slow setup, shared accounts, or hard-coded data in the seed spread to every generated file. Fixing the seed after generating fifty tests means regenerating or editing all fifty.

Definitions go stale. The agent files are tied to the Playwright version that wrote them. Upgrading Playwright without rerunning init-agents leaves the agents calling tools that changed.

Planner plans vs written test cases

The planner's Markdown plan looks a lot like a manual test case: steps, expected results, data. The difference is where the expected results come from. A planner derives them from the app as it behaves today. A test case written from the requirement derives them from what the app is supposed to do. The second kind catches the bug on day one; the first kind catches the next change after it.

A practical split: write or generate test cases from the spec first, drop them into specs/ in the planner's Markdown format, and let the generator turn them into code. You keep the planner for exploring screens nobody wrote a spec for.

What is worth automating

The agents handle the expensive mechanical parts of end-to-end work well: finding stable locators, writing the boilerplate around each step, and repairing tests after a harmless UI change. Deciding what the expected result should be, and approving a healer patch that changes one, stays with a person who knows the requirement. If you are deciding how much of the browser work to give an agent in the first place, the Playwright MCP guide and the CLI vs MCP comparison cover the two ways it can drive the browser.

Open sourceThe skill lives on GitHubMIT licensed, no account, nothing to install on your machine. Star it if it saves you an afternoon.View the repository