Playwright MCP vs Playwright CLI: Browser Testing for Coding Agents
Playwright MCP (@playwright/mcp) and Playwright CLI (@playwright/cli) are Microsoft’s two ways to let a coding agent drive a real browser through Playwright’s accessibility snapshots. The MCP server keeps a live browser and returns page state inside every tool result. The CLI with its skill costs less context, and Microsoft’s README points coding agents to it.
You ask the agent to fix a broken Subscribe button. It edits the click handler, runs the unit tests, and reports success, and nobody opened a browser. Or it did open one, and three agents in three worktrees fought over the same browser profile until two of them failed with a locked-profile error.
This page is for developers who want the agent to check its own UI work in a real browser, and for tech leads who need that to work with several agents running at once.
What you get from this page
Section titled “What you get from this page”- A decision rule for Playwright MCP or Playwright CLI plus its skill, based on measured context cost.
- Install commands for Claude Code, Codex and Cursor, with the agent flags
--isolated,--headlessand--caps. - A bug-report-to-regression-test workflow with its CI gate: the test fails first, the fix makes it pass, and CI keeps it.
Should your agent use Playwright MCP or Playwright CLI?
Section titled “Should your agent use Playwright MCP or Playwright CLI?”Both packages drive the same Playwright engine and use the same element references from the accessibility snapshot. They differ in where the page data goes.
- Playwright MCP is a stdio MCP server. Each action (
browser_click,browser_navigate) returns the resulting page snapshot inside the tool result, so the page lands in the model’s context after every step. - Playwright CLI is a shell command,
playwright-cli. It writes the snapshot to a YAML file under.playwright-cli/and prints a link to it. The agent reads the file only when it needs to.
The Playwright MCP README (0.0.82, read 2026-09-26) states the choice outright: “If you are using a coding agent, you might benefit from using the CLI+SKILLS instead.” Its reason is that CLI calls “avoid loading large tool schemas and verbose accessibility trees into the model context.” It keeps MCP for “exploratory automation, self-healing tests, or long-running autonomous workflows where maintaining continuous browser context outweighs token cost concerns.”
| Playwright MCP | Playwright CLI + skill | |
|---|---|---|
| Package (npm, 2026-09-26) | @playwright/mcp 0.0.82 | @playwright/cli 0.1.21, binary playwright-cli |
| How the agent calls it | MCP tools (browser_navigate, browser_snapshot, …) | Shell commands (playwright-cli goto, playwright-cli snapshot, …) |
| Page state after each action | Inline in the tool result | Written to a file; the agent reads it on demand |
| Browser profile by default | Persistent, on disk, one per workspace | In memory; --persistent saves it |
| Works in | Any MCP client | Any agent that can run shell commands; the skill needs an agent that loads skills |
| Best for | Exploratory sessions, a human watching a headed browser, clients without a shell | Coding agents in a repository: test writing, verification, parallel sessions |
The decision rule: default to the CLI plus skill in a repository where the agent can run shell commands. Keep the MCP server when the agent has no shell, when you explore an unfamiliar app interactively, or when the team already manages MCP servers centrally and the context budget is not the constraint.
Install Playwright MCP in Claude Code, Codex and Cursor
Section titled “Install Playwright MCP in Claude Code, Codex and Cursor”The server needs Node.js 18 or newer. Add --isolated from the start if more than one agent can run in the same checkout; the section on parallel agents explains why.
Project scope, so every clone and worktree gets the same entry (tested on Claude Code 2.1.283):
claude mcp add --scope project playwright -- npx @playwright/mcp@0.0.82 --isolated --headlessThat writes this entry to .mcp.json, which you commit:
{ "mcpServers": { "playwright": { "type": "stdio", "command": "npx", "args": ["@playwright/mcp@0.0.82", "--isolated", "--headless"], "env": {} } }}The -- matters: without it, claude mcp add tries to parse --isolated as its own option. The README’s shorter line, claude mcp add playwright npx @playwright/mcp@latest, works when you pass no server flags.
Alternatively, install the plugin from Anthropic’s official marketplace. Its .mcp.json runs npx @playwright/mcp@latest with no flags, so it uses the persistent profile:
claude plugin install playwright@claude-plugins-officialTested on codex-cli 0.157.1:
codex mcp add playwright -- npx @playwright/mcp@0.0.82 --isolated --headlessThat writes a global entry to ~/.codex/config.toml:
[mcp_servers.playwright]command = "npx"args = ["@playwright/mcp@0.0.82", "--isolated", "--headless"]Because the entry is global, it applies to every Codex session on the machine, including sessions started with codex --worktree. That is a reason to put --isolated in it rather than leaving it out. The README’s own line is codex mcp add playwright npx "@playwright/mcp@latest".
Add the server in Cursor Settings > MCP > Add new MCP Server as a command type with npx @playwright/mcp@latest, or put the standard configuration in .cursor/mcp.json (shape from the Playwright MCP README; not tested in a Cursor binary):
{ "mcpServers": { "playwright": { "command": "npx", "args": ["@playwright/mcp@0.0.82", "--isolated", "--headless"] } }}Confirm it loaded: run /mcp in Claude Code or Codex, or codex mcp list in a terminal, and check that playwright is connected. In Cursor, open the MCP list in settings.
Install Playwright CLI and its skill
Section titled “Install Playwright CLI and its skill”Install the CLI once per machine, then install the skill into each repository so it travels with the code (tested with @playwright/cli 0.1.21):
npm install -g @playwright/cli@latestplaywright-cli --helpplaywright-cli install --skillsThe default target is claude: the command writes .claude/skills/playwright-cli/SKILL.md plus a references/ folder with guides for test generation, running tests, request mocking, storage state and tracing. Commit the folder. Add --global to install into your home directory instead.
playwright-cli install --skills agentsThis writes the same skill to .agents/skills/playwright-cli/, the repository folder Codex scans for shared skills (see Codex team workflows). Run /skills in a Codex session to confirm it is listed.
playwright-cli install --skills agentsThis writes the skill to .agents/skills/playwright-cli/. Cursor’s skill folders could not be checked against cursor.com on 2026-09-26, so confirm in Cursor that the skill appears. If it does not, use the README’s skills-less route: tell the agent to run playwright-cli --help and work from the command list.
The skill’s frontmatter pre-approves Bash(playwright-cli:*), Bash(npx:*) and Bash(npm:*) in agents that honour allowed-tools. Read that line before you commit the skill: npx:* and npm:* are broader than browser testing needs.
Flags that matter when an agent drives Playwright MCP
Section titled “Flags that matter when an agent drives Playwright MCP”All flags below are from playwright-mcp --help and the README, version 0.0.82. Each also has an environment variable, such as PLAYWRIGHT_MCP_ISOLATED.
| Flag | What it does | When to set it |
|---|---|---|
--isolated | Keeps the browser profile in memory and does not save it to disk | Always when more than one agent can run; always in CI |
--headless | Runs without a visible window (headed is the default) | Background agents, CI, remote containers |
--caps testing | Adds browser_generate_locator and the browser_verify_* assertion tools | When the agent writes tests from what it sees |
--caps network | Adds browser_route, browser_unroute and browser_network_state_set for mocking and offline mode | Reproducing a bug that depends on an API response |
--caps vision,pdf,devtools | Coordinate-based mouse tools, PDF export, tracing and video | Canvas-heavy UIs, PDF output, trace capture |
--storage-state <path> | Loads cookies and local storage into an isolated context | Testing behind a login without a persistent profile |
--user-data-dir <path> | Uses a specific persistent profile directory | When you need a persistent profile per agent |
--console-level error | Returns only console messages at that level and above | Keeping console output short |
--output-dir <path> | Where automatically named screenshots and files go | Collecting evidence for the pull request |
Check a page: open pricing, click Subscribe, report errors
Section titled “Check a page: open pricing, click Subscribe, report errors”Run this first: it answers what a unit test cannot, whether the page works in a browser. Start your dev server first. The prompt works in all three tools.
What you should see with Playwright MCP: browser_navigate, browser_snapshot, browser_click (with an element reference from the snapshot, such as e21), then browser_console_messages with level: "error", browser_network_requests, and browser_take_screenshot. browser_network_requests leaves out successful static resources by default, which keeps the list short.
What you should see with Playwright CLI: the same steps as shell commands.
playwright-cli open http://localhost:4321/en/pricingplaywright-cli snapshotplaywright-cli click e21playwright-cli console errorplaywright-cli requestsplaywright-cli screenshotIf the agent reports “no errors” without having called the console and network tools, it guessed. Ask it to show the tool output.
Turn a bug report into a regression test with Playwright
Section titled “Turn a bug report into a regression test with Playwright”The smoke check proves the page works today. A test proves it keeps working. The workflow below is the one to standardise on: the agent reproduces the bug in a browser, writes a Playwright test that fails for the right reason, fixes the code, and proves the same test passes. It uses @playwright/test (1.63.0 on npm, 2026-09-26) as the runner, whichever driver the agent explores with.
-
Give the agent the bug report and forbid the fix. The first deliverable is a failing test, not a patch.
-
Check that it fails for the right reason. The failure must be the bug from the report (the missing checkout request or the
TypeError), not a timeout on a wrong locator or a server that was not running. A test that fails for the wrong reason proves nothing when it later passes. Once it fails for the right reason, commit the test alone (test: reproduce pricing subscribe bug), so the red state has its own commit. -
Let the agent fix the code, not the test.
--repeat-each=5runs the test five times. A test that passes four times out of five is a flaky test, not a fix. -
Keep the test in CI. The pull request holds two commits: the red test, then the fix. The CI job runs the E2E suite with
--fail-on-flaky-tests, so a test that only passes on retry fails the build, and it checks that the spec did not change after the red commit:# .github/workflows/e2e.yml (job steps; checkout needs fetch-depth: 0)- run: npx playwright install --with-deps chromium# --retries=2 lets Playwright detect a flaky test; --fail-on-flaky-tests fails the build on one- run: npx playwright test --retries=2 --fail-on-flaky-tests- name: Spec unchanged since the red commitrun: |RED_SHA=$(git log --format=%H --grep='^test: reproduce pricing subscribe bug' -n 1)git diff --exit-code "$RED_SHA" HEAD -- tests/e2e/pricing-subscribe.spec.tsThe diff step fails the build if the fix commit touched the test.
-
Attach the evidence. The pull request carries the red run output from step 1, the green runs from step 3, the trace, and the screenshot. The evidence bundle page has the template that makes these fields required.
With Playwright CLI, the skill adds a shortcut to step 1: each playwright-cli action prints the equivalent Playwright TypeScript (await page.getByRole('button', { name: 'Subscribe' }).click();), so the agent assembles the test from the actions it already took. The skill’s references/test-generation.md calls this plan, generate and heal. With Playwright MCP, --caps testing gives the agent browser_generate_locator for the same purpose.
Run one agent per worktree without browser collisions
Section titled “Run one agent per worktree without browser collisions”The Playwright MCP README is explicit: “A persistent profile can only be used by one browser instance at a time, so concurrent MCP clients sharing the same workspace will conflict.” The fix it gives is to start each additional client with --isolated or with a distinct --user-data-dir.
The default profile path ends in {workspace-hash}, derived from the client’s workspace root, so two worktrees usually get separate profiles already. Two agents in the same checkout do not, and the hash depends on what the client reports as its root. --isolated removes the question: every session starts from a clean in-memory profile, which is also what a test should start from.
Browser profiles are only half of it. The other shared resource is the dev server port:
| Shared resource | What goes wrong | Setting that prevents it |
|---|---|---|
| Playwright MCP profile | Second browser fails to start, or reuses another agent’s login | --isolated in the committed MCP entry |
| Playwright CLI session | Two agents drive the same named browser | Start each agent with its own PLAYWRIGHT_CLI_SESSION, or prepend -s=<name> to each command |
| Dev-server port | Astro, Vite and Next.js move to the next free port without failing; the agent tests another worktree’s server | A fixed port block per worktree |
reuseExistingServer in playwright.config.ts | Playwright attaches to whatever already listens on the port and reports green against code that was never under test | Derive baseURL, webServer.url and the dev-server port from one per-worktree value |
For Playwright CLI, the README’s pattern is an environment variable at agent start. In a worktree, name the session after the directory:
PLAYWRIGHT_CLI_SESSION="$(basename "$PWD")" claudeRun playwright-cli list to see every open session, and playwright-cli show to open a dashboard with a live view of each one. For the full worktree setup, including port blocks, see running many agents at once.
How to prove the browser check means something
Section titled “How to prove the browser check means something”A browser check makes the agent’s claim observable. It proves the change only when the check could have failed and is kept where the agent cannot quietly weaken it.
| Check | What it proves | Who owns it |
|---|---|---|
| Red run before the fix, with the failure message | The test detects this bug | Author (agent); the reviewer reads the message |
Green runs with --repeat-each=5 --retries=0 | The fix works and the test is not flaky | Author (agent); CI reruns it |
| Test file unchanged between the red and green commits | The agent fixed the code, not the test | CI diff check; see protecting the oracle |
E2E suite in CI with --fail-on-flaky-tests or zero retries | The behaviour stays fixed on every later change | CI, blocking |
| Trace and screenshot attached to the pull request | A reviewer can replay what the browser did without rerunning it | Author (agent) |
The tech lead approves on that evidence. For the reviewing side of the loop, see reviewing an agent’s pull request.
What Playwright MCP and Playwright CLI cost in context
Section titled “What Playwright MCP and Playwright CLI cost in context”Measured on 2026-09-26 by listing the tools of @playwright/mcp 0.0.82 over stdio:
| Configuration | Tools | Tool-schema JSON |
|---|---|---|
| Default | 25 | about 20,100 characters |
--caps testing,network | 34 | about 25,700 characters |
| All seven capability groups | 72 | about 45,400 characters |
| Playwright CLI skill | 0 (shell commands) | the SKILL.md body is 15,179 characters, loaded when the skill triggers |
The schemas are the smaller cost. Claude Code 2.1.283 and codex-cli 0.157.1 both load MCP tool schemas on demand through tool search. The larger cost is the accessibility snapshot that Playwright MCP returns after every action: it grows with the size of the page and repeats with every click, so a ten-step flow on a content-heavy page pays for ten snapshots.
Ways to cut it, cheapest first:
- Use Playwright CLI, which writes snapshots to files.
- On Playwright MCP, ask for
browser_snapshotwithdepthor atargetelement, or withfilenameto save it instead of returning it. - Start the server with
--snapshot-mode nonewhen the agent drives a known flow and does not need the page after each step. - Add only the
--capsgroups the task needs. - Push long browser sessions into a subagent so only its summary returns to the main session.
To measure it yourself in Claude Code, run /context before and after the pricing-page check, once with each driver. For the general pattern, see reducing MCP token cost.
How popular are Playwright MCP and Playwright CLI?
Section titled “How popular are Playwright MCP and Playwright CLI?”As of 2026-09-26: microsoft/playwright-mcp has 37.6k GitHub stars and is listed in the Official MCP Registry as io.github.microsoft/playwright-mcp 0.0.82; the playwright plugin has 319,887 installs on the Claude plugin directory (claude.com/plugins). microsoft/playwright-cli has 13.6k GitHub stars, and its skill shows 165,672 all-time installs in a third-party scrape of skills.sh (LinklyAI/best-skills, dated 2026-09-26; secondary, because skills.sh itself could not be read).
When Playwright MCP or Playwright CLI breaks
Section titled “When Playwright MCP or Playwright CLI breaks”The browser fails to launch or reports a locked profile. Another agent or an earlier session holds the persistent profile. Recovery: add --isolated to the MCP entry, or give each agent its own --user-data-dir. For a stuck CLI session, run playwright-cli list, then playwright-cli close-all, and playwright-cli kill-all if processes remain.
No browser is installed. A fresh machine or CI container has no browser binaries, and the first navigation fails. Recovery: run npx playwright install chromium (add --with-deps on Linux CI images), or playwright-cli install-browser for the CLI. Choose a channel explicitly with --browser chrome|firefox|webkit|msedge.
The test passed against the wrong server. The dev server moved to the next free port, and the agent or reuseExistingServer hit another worktree’s server. Recovery: pin a port per worktree, derive baseURL and webServer.url from it, and ask the agent to prove which server it tested, as the worktree prompt does.
The agent edits the test until it passes. The red run was real; the green run came from a weaker assertion. Recovery: make “test file unchanged” a CI check, and put tests/e2e/ under a deny rule or CODEOWNERS as described in protecting the oracle.
The test is flaky. It passes locally and fails in CI, often on a fixed waitForTimeout or a CSS selector. Recovery: rerun with --repeat-each=10 --retries=0 and --trace=on, open the trace, and ask the agent to replace sleeps with web-first assertions and CSS selectors with role-based locators.
The login state is gone. --isolated discards cookies when the browser closes. Recovery: save a storage state once with Playwright’s setup project or playwright-cli state-save auth.json. On Playwright MCP, pass it with --storage-state auth.json; on Playwright CLI, run playwright-cli state-load auth.json. Keep that file out of Git; it holds session cookies.
The context window fills after a few clicks. Each MCP action returned a full snapshot of a large page. Recovery: switch that task to Playwright CLI, or use the snapshot limits in the previous section.