Skip to content

Playwright CLI vs MCP: Which Should Your Agent Use?

ILIA KARPENKO5 min read

Playwright CLI and Playwright MCP are two interfaces Microsoft ships for letting an AI agent control a browser through Playwright. Playwright MCP is a Model Context Protocol server: the agent calls its tools and gets the page state back inside the conversation. playwright-cli is a command-line tool: the agent runs shell commands, and the page state goes to files on disk that the agent reads only when it needs them. For coding agents like Claude Code and GitHub Copilot, the Playwright team recommends the CLI. For long exploratory loops that reason over the page step by step, MCP still fits better.

The difference shows up in the bill and in how long a session lasts. An MCP server returns a full accessibility snapshot after each action, and those snapshots sit in the model's context for the rest of the session. On a dense app, a twenty-step flow can fill a large share of the context window before the agent has written a line of test code. The CLI writes the same snapshot to a YAML file and returns a short status, so the agent pays for the page only when it decides to look.

Which "Playwright CLI" do you mean?

The phrase covers two different tools, and search results mix them up.

ToolCommandWhat it is for
Playwright Test CLInpx playwright test, codegen, show-reportRunning and debugging the test suite you already have
playwright-cliplaywright-cli open, click, snapshotLetting a coding agent drive a live browser, one command at a time

The first ships with @playwright/test and has been around since the start: npx playwright test --ui opens UI mode, npx playwright codegen https://example.com records your clicks as code, npx playwright show-report opens the HTML report. The second is the newer @playwright/cli package, built for agents. The rest of this post is about the second one and how it compares to the MCP server.

Free and open sourceCasely writes these cases for youAttach a spec and one file of your team's existing test cases in Claude. Casely copies your columns, names the gaps it found in the spec, and exports a single Excel file your tracker imports in one pass.Read the install docs

How to install playwright-cli

You need Node.js 20 or newer. Install it globally, then add the skills so the agent knows the commands without reading the help text every time:

npm install -g @playwright/cli@latest
playwright-cli install --skills

The skills land in the project's .claude/skills folder, where Claude Code picks them up. Add -g to install them into your home directory for every project. Skills are optional: an agent told "use playwright-cli, check playwright-cli --help" will work out the commands itself.

A session looks like this when you run it by hand:

playwright-cli open https://demo.playwright.dev/todomvc/ --headed
playwright-cli type "Buy groceries"
playwright-cli press Enter
playwright-cli snapshot
playwright-cli check e21
playwright-cli screenshot

Each command prints the page URL and title, plus a path to the snapshot file. Element references like e21 come from that snapshot, the same way they do in MCP. You can also target elements with CSS or a Playwright locator such as "getByRole('button', { name: 'Submit' })".

How the two compare

playwright-cliPlaywright MCP
How the agent calls itShell commandsMCP tool calls
Page snapshotWritten to a file, read on demandReturned into the model's context after every action
Tool schema in contextNone, or a short skill fileFull schema for every enabled tool
Setup in Claude Codenpm install -g @playwright/cli, then skillsclaude mcp add playwright npx @playwright/mcp@latest
Works without shell accessNoYes
Browser stateIn memory per session; --persistent saves it to diskPersistent profile by default; --isolated for clean runs
Several browsers at onceNamed sessions with -s=nameOne server per browser, or isolated contexts
Watching what the agent doesplaywright-cli show opens a live dashboardHeaded browser window
Recording actions as coderecording-start and recording-stopbrowser_start_recording with the devtools capability

Both expose the same underlying abilities: navigation, form input, network mocking, cookies and storage, console and network inspection, tracing, video, and locator generation. The choice is about the interface, not about what the browser can do.

When to use playwright-cli

Pick the CLI when the agent already has a terminal and the browser is one part of a bigger job. Writing a feature, running its tests, opening the page to check the result, and fixing what broke all compete for the same context window. Keeping snapshots on disk leaves room for the code.

The named sessions help here too. playwright-cli -s=admin open and playwright-cli -s=customer open give an agent two logged-in browsers for a test that needs both roles. Setting PLAYWRIGHT_CLI_SESSION=checkout before starting the agent keeps it on one session without repeating the flag.

When Playwright MCP still wins

Pick MCP when the agent has no shell, as in Claude Desktop or a hosted chat client, or when the task is the browser: exploring an unfamiliar app, reasoning over the page structure after every step, or running a long autonomous loop where keeping the full page in view matters more than tokens. The Playwright MCP guide covers its tools, capability flags, and setup in detail.

Some teams keep both. The cost of having MCP configured is the tool schema in every session, so it is cheaper to add it to a project only when someone needs it than to leave it on everywhere.

Where both fall short

Neither tool decides what to test. An agent that clicks through a checkout and finds no error has proven the checkout did not crash, which is a weaker claim than "the checkout charges the right amount." The expected results have to come from somewhere outside the browser: a requirement, an acceptance criterion, a written test case.

Neither one is a security boundary either. Both drive a real browser with whatever cookies and files you give them. Use test accounts, keep production credentials out of the profile, and treat any text the page shows the agent as untrusted input, not as instructions.

The CLI also has a trap of its own. Because snapshots go to disk, an agent can act on a reference from a snapshot that is several actions old. If clicks start landing on the wrong element, have it take a fresh playwright-cli snapshot before acting.

What is worth automating

Driving the browser is the part an agent does well with either tool: reproducing a bug, checking a fix in the real page, turning a manual walk into a draft spec with stable locators. Writing the expected results is the part to keep upstream. A test case with a concrete expected result gives the agent something to verify, so it does not invent a check from whatever the page happened to show.

Open sourceThe skill lives on GitHubMIT licensed, no account, nothing to install on your machine. Star it if it saves you an afternoon.View the repository