Playwright CLI vs MCP: Which Should Your Agent Use?
Playwright CLI and Playwright MCP are two interfaces Microsoft ships for letting an AI agent control a browser through Playwright. Playwright MCP is a Model Context Protocol server: the agent calls its tools and gets the page state back inside the conversation. playwright-cli is a command-line tool: the agent runs shell commands, and the page state goes to files on disk that the agent reads only when it needs them. For coding agents like Claude Code and GitHub Copilot, the Playwright team recommends the CLI. For long exploratory loops that reason over the page step by step, MCP still fits better.
The difference shows up in the bill and in how long a session lasts. An MCP server returns a full accessibility snapshot after each action, and those snapshots sit in the model's context for the rest of the session. On a dense app, a twenty-step flow can fill a large share of the context window before the agent has written a line of test code. The CLI writes the same snapshot to a YAML file and returns a short status, so the agent pays for the page only when it decides to look.
Which "Playwright CLI" do you mean?
The phrase covers two different tools, and search results mix them up.
| Tool | Command | What it is for |
|---|---|---|
| Playwright Test CLI | npx playwright test, codegen, show-report | Running and debugging the test suite you already have |
| playwright-cli | playwright-cli open, click, snapshot | Letting a coding agent drive a live browser, one command at a time |
The first ships with @playwright/test and has been around since the start: npx playwright test --ui opens UI mode, npx playwright codegen https://example.com records your clicks as code, npx playwright show-report opens the HTML report. The second is the newer @playwright/cli package, built for agents. The rest of this post is about the second one and how it compares to the MCP server.
How to install playwright-cli
You need Node.js 20 or newer. Install it globally, then add the skills so the agent knows the commands without reading the help text every time:
npm install -g @playwright/cli@latest
playwright-cli install --skills
The skills land in the project's .claude/skills folder, where Claude Code picks them up. Add -g to install them into your home directory for every project. Skills are optional: an agent told "use playwright-cli, check playwright-cli --help" will work out the commands itself.
A session looks like this when you run it by hand:
playwright-cli open https://demo.playwright.dev/todomvc/ --headed
playwright-cli type "Buy groceries"
playwright-cli press Enter
playwright-cli snapshot
playwright-cli check e21
playwright-cli screenshot
Each command prints the page URL and title, plus a path to the snapshot file. Element references like e21 come from that snapshot, the same way they do in MCP. You can also target elements with CSS or a Playwright locator such as "getByRole('button', { name: 'Submit' })".
How the two compare
| playwright-cli | Playwright MCP | |
|---|---|---|
| How the agent calls it | Shell commands | MCP tool calls |
| Page snapshot | Written to a file, read on demand | Returned into the model's context after every action |
| Tool schema in context | None, or a short skill file | Full schema for every enabled tool |
| Setup in Claude Code | npm install -g @playwright/cli, then skills | claude mcp add playwright npx @playwright/mcp@latest |
| Works without shell access | No | Yes |
| Browser state | In memory per session; --persistent saves it to disk | Persistent profile by default; --isolated for clean runs |
| Several browsers at once | Named sessions with -s=name | One server per browser, or isolated contexts |
| Watching what the agent does | playwright-cli show opens a live dashboard | Headed browser window |
| Recording actions as code | recording-start and recording-stop | browser_start_recording with the devtools capability |
Both expose the same underlying abilities: navigation, form input, network mocking, cookies and storage, console and network inspection, tracing, video, and locator generation. The choice is about the interface, not about what the browser can do.
When to use playwright-cli
Pick the CLI when the agent already has a terminal and the browser is one part of a bigger job. Writing a feature, running its tests, opening the page to check the result, and fixing what broke all compete for the same context window. Keeping snapshots on disk leaves room for the code.
The named sessions help here too. playwright-cli -s=admin open and playwright-cli -s=customer open give an agent two logged-in browsers for a test that needs both roles. Setting PLAYWRIGHT_CLI_SESSION=checkout before starting the agent keeps it on one session without repeating the flag.
When Playwright MCP still wins
Pick MCP when the agent has no shell, as in Claude Desktop or a hosted chat client, or when the task is the browser: exploring an unfamiliar app, reasoning over the page structure after every step, or running a long autonomous loop where keeping the full page in view matters more than tokens. The Playwright MCP guide covers its tools, capability flags, and setup in detail.
Some teams keep both. The cost of having MCP configured is the tool schema in every session, so it is cheaper to add it to a project only when someone needs it than to leave it on everywhere.
Where both fall short
Neither tool decides what to test. An agent that clicks through a checkout and finds no error has proven the checkout did not crash, which is a weaker claim than "the checkout charges the right amount." The expected results have to come from somewhere outside the browser: a requirement, an acceptance criterion, a written test case.
Neither one is a security boundary either. Both drive a real browser with whatever cookies and files you give them. Use test accounts, keep production credentials out of the profile, and treat any text the page shows the agent as untrusted input, not as instructions.
The CLI also has a trap of its own. Because snapshots go to disk, an agent can act on a reference from a snapshot that is several actions old. If clicks start landing on the wrong element, have it take a fresh playwright-cli snapshot before acting.
What is worth automating
Driving the browser is the part an agent does well with either tool: reproducing a bug, checking a fix in the real page, turning a manual walk into a draft spec with stable locators. Writing the expected results is the part to keep upstream. A test case with a concrete expected result gives the agent something to verify, so it does not invent a check from whatever the page happened to show.