Skip to content

Playwright MCP vs Playwright CLI: Browser Testing for Coding Agents

Playwright MCP (@playwright/mcp) and Playwright CLI (@playwright/cli) are Microsoft’s two ways to let a coding agent drive a real browser through Playwright’s accessibility snapshots. The MCP server keeps a live browser and returns page state inside every tool result. The CLI with its skill costs less context, and Microsoft’s README points coding agents to it.

You ask the agent to fix a broken Subscribe button. It edits the click handler, runs the unit tests, and reports success, and nobody opened a browser. Or it did open one, and three agents in three worktrees fought over the same browser profile until two of them failed with a locked-profile error.

This page is for developers who want the agent to check its own UI work in a real browser, and for tech leads who need that to work with several agents running at once.

  • A decision rule for Playwright MCP or Playwright CLI plus its skill, based on measured context cost.
  • Install commands for Claude Code, Codex and Cursor, with the agent flags --isolated, --headless and --caps.
  • A bug-report-to-regression-test workflow with its CI gate: the test fails first, the fix makes it pass, and CI keeps it.

Should your agent use Playwright MCP or Playwright CLI?

Section titled “Should your agent use Playwright MCP or Playwright CLI?”

Both packages drive the same Playwright engine and use the same element references from the accessibility snapshot. They differ in where the page data goes.

  • Playwright MCP is a stdio MCP server. Each action (browser_click, browser_navigate) returns the resulting page snapshot inside the tool result, so the page lands in the model’s context after every step.
  • Playwright CLI is a shell command, playwright-cli. It writes the snapshot to a YAML file under .playwright-cli/ and prints a link to it. The agent reads the file only when it needs to.

The Playwright MCP README (0.0.82, read 2026-09-26) states the choice outright: “If you are using a coding agent, you might benefit from using the CLI+SKILLS instead.” Its reason is that CLI calls “avoid loading large tool schemas and verbose accessibility trees into the model context.” It keeps MCP for “exploratory automation, self-healing tests, or long-running autonomous workflows where maintaining continuous browser context outweighs token cost concerns.”

Playwright MCPPlaywright CLI + skill
Package (npm, 2026-09-26)@playwright/mcp 0.0.82@playwright/cli 0.1.21, binary playwright-cli
How the agent calls itMCP tools (browser_navigate, browser_snapshot, …)Shell commands (playwright-cli goto, playwright-cli snapshot, …)
Page state after each actionInline in the tool resultWritten to a file; the agent reads it on demand
Browser profile by defaultPersistent, on disk, one per workspaceIn memory; --persistent saves it
Works inAny MCP clientAny agent that can run shell commands; the skill needs an agent that loads skills
Best forExploratory sessions, a human watching a headed browser, clients without a shellCoding agents in a repository: test writing, verification, parallel sessions

The decision rule: default to the CLI plus skill in a repository where the agent can run shell commands. Keep the MCP server when the agent has no shell, when you explore an unfamiliar app interactively, or when the team already manages MCP servers centrally and the context budget is not the constraint.

Install Playwright MCP in Claude Code, Codex and Cursor

Section titled “Install Playwright MCP in Claude Code, Codex and Cursor”

The server needs Node.js 18 or newer. Add --isolated from the start if more than one agent can run in the same checkout; the section on parallel agents explains why.

Project scope, so every clone and worktree gets the same entry (tested on Claude Code 2.1.283):

Terminal window
claude mcp add --scope project playwright -- npx @playwright/mcp@0.0.82 --isolated --headless

That writes this entry to .mcp.json, which you commit:

{
"mcpServers": {
"playwright": {
"type": "stdio",
"command": "npx",
"args": ["@playwright/mcp@0.0.82", "--isolated", "--headless"],
"env": {}
}
}
}

The -- matters: without it, claude mcp add tries to parse --isolated as its own option. The README’s shorter line, claude mcp add playwright npx @playwright/mcp@latest, works when you pass no server flags.

Alternatively, install the plugin from Anthropic’s official marketplace. Its .mcp.json runs npx @playwright/mcp@latest with no flags, so it uses the persistent profile:

Terminal window
claude plugin install playwright@claude-plugins-official

Confirm it loaded: run /mcp in Claude Code or Codex, or codex mcp list in a terminal, and check that playwright is connected. In Cursor, open the MCP list in settings.

Install the CLI once per machine, then install the skill into each repository so it travels with the code (tested with @playwright/cli 0.1.21):

Terminal window
npm install -g @playwright/cli@latest
playwright-cli --help
Terminal window
playwright-cli install --skills

The default target is claude: the command writes .claude/skills/playwright-cli/SKILL.md plus a references/ folder with guides for test generation, running tests, request mocking, storage state and tracing. Commit the folder. Add --global to install into your home directory instead.

The skill’s frontmatter pre-approves Bash(playwright-cli:*), Bash(npx:*) and Bash(npm:*) in agents that honour allowed-tools. Read that line before you commit the skill: npx:* and npm:* are broader than browser testing needs.

Flags that matter when an agent drives Playwright MCP

Section titled “Flags that matter when an agent drives Playwright MCP”

All flags below are from playwright-mcp --help and the README, version 0.0.82. Each also has an environment variable, such as PLAYWRIGHT_MCP_ISOLATED.

FlagWhat it doesWhen to set it
--isolatedKeeps the browser profile in memory and does not save it to diskAlways when more than one agent can run; always in CI
--headlessRuns without a visible window (headed is the default)Background agents, CI, remote containers
--caps testingAdds browser_generate_locator and the browser_verify_* assertion toolsWhen the agent writes tests from what it sees
--caps networkAdds browser_route, browser_unroute and browser_network_state_set for mocking and offline modeReproducing a bug that depends on an API response
--caps vision,pdf,devtoolsCoordinate-based mouse tools, PDF export, tracing and videoCanvas-heavy UIs, PDF output, trace capture
--storage-state <path>Loads cookies and local storage into an isolated contextTesting behind a login without a persistent profile
--user-data-dir <path>Uses a specific persistent profile directoryWhen you need a persistent profile per agent
--console-level errorReturns only console messages at that level and aboveKeeping console output short
--output-dir <path>Where automatically named screenshots and files goCollecting evidence for the pull request

Check a page: open pricing, click Subscribe, report errors

Section titled “Check a page: open pricing, click Subscribe, report errors”

Run this first: it answers what a unit test cannot, whether the page works in a browser. Start your dev server first. The prompt works in all three tools.

What you should see with Playwright MCP: browser_navigate, browser_snapshot, browser_click (with an element reference from the snapshot, such as e21), then browser_console_messages with level: "error", browser_network_requests, and browser_take_screenshot. browser_network_requests leaves out successful static resources by default, which keeps the list short.

What you should see with Playwright CLI: the same steps as shell commands.

Terminal window
playwright-cli open http://localhost:4321/en/pricing
playwright-cli snapshot
playwright-cli click e21
playwright-cli console error
playwright-cli requests
playwright-cli screenshot

If the agent reports “no errors” without having called the console and network tools, it guessed. Ask it to show the tool output.

Turn a bug report into a regression test with Playwright

Section titled “Turn a bug report into a regression test with Playwright”

The smoke check proves the page works today. A test proves it keeps working. The workflow below is the one to standardise on: the agent reproduces the bug in a browser, writes a Playwright test that fails for the right reason, fixes the code, and proves the same test passes. It uses @playwright/test (1.63.0 on npm, 2026-09-26) as the runner, whichever driver the agent explores with.

  1. Give the agent the bug report and forbid the fix. The first deliverable is a failing test, not a patch.

  2. Check that it fails for the right reason. The failure must be the bug from the report (the missing checkout request or the TypeError), not a timeout on a wrong locator or a server that was not running. A test that fails for the wrong reason proves nothing when it later passes. Once it fails for the right reason, commit the test alone (test: reproduce pricing subscribe bug), so the red state has its own commit.

  3. Let the agent fix the code, not the test.

    --repeat-each=5 runs the test five times. A test that passes four times out of five is a flaky test, not a fix.

  4. Keep the test in CI. The pull request holds two commits: the red test, then the fix. The CI job runs the E2E suite with --fail-on-flaky-tests, so a test that only passes on retry fails the build, and it checks that the spec did not change after the red commit:

    # .github/workflows/e2e.yml (job steps; checkout needs fetch-depth: 0)
    - run: npx playwright install --with-deps chromium
    # --retries=2 lets Playwright detect a flaky test; --fail-on-flaky-tests fails the build on one
    - run: npx playwright test --retries=2 --fail-on-flaky-tests
    - name: Spec unchanged since the red commit
    run: |
    RED_SHA=$(git log --format=%H --grep='^test: reproduce pricing subscribe bug' -n 1)
    git diff --exit-code "$RED_SHA" HEAD -- tests/e2e/pricing-subscribe.spec.ts

    The diff step fails the build if the fix commit touched the test.

  5. Attach the evidence. The pull request carries the red run output from step 1, the green runs from step 3, the trace, and the screenshot. The evidence bundle page has the template that makes these fields required.

With Playwright CLI, the skill adds a shortcut to step 1: each playwright-cli action prints the equivalent Playwright TypeScript (await page.getByRole('button', { name: 'Subscribe' }).click();), so the agent assembles the test from the actions it already took. The skill’s references/test-generation.md calls this plan, generate and heal. With Playwright MCP, --caps testing gives the agent browser_generate_locator for the same purpose.

Run one agent per worktree without browser collisions

Section titled “Run one agent per worktree without browser collisions”

The Playwright MCP README is explicit: “A persistent profile can only be used by one browser instance at a time, so concurrent MCP clients sharing the same workspace will conflict.” The fix it gives is to start each additional client with --isolated or with a distinct --user-data-dir.

The default profile path ends in {workspace-hash}, derived from the client’s workspace root, so two worktrees usually get separate profiles already. Two agents in the same checkout do not, and the hash depends on what the client reports as its root. --isolated removes the question: every session starts from a clean in-memory profile, which is also what a test should start from.

Browser profiles are only half of it. The other shared resource is the dev server port:

Shared resourceWhat goes wrongSetting that prevents it
Playwright MCP profileSecond browser fails to start, or reuses another agent’s login--isolated in the committed MCP entry
Playwright CLI sessionTwo agents drive the same named browserStart each agent with its own PLAYWRIGHT_CLI_SESSION, or prepend -s=<name> to each command
Dev-server portAstro, Vite and Next.js move to the next free port without failing; the agent tests another worktree’s serverA fixed port block per worktree
reuseExistingServer in playwright.config.tsPlaywright attaches to whatever already listens on the port and reports green against code that was never under testDerive baseURL, webServer.url and the dev-server port from one per-worktree value

For Playwright CLI, the README’s pattern is an environment variable at agent start. In a worktree, name the session after the directory:

Terminal window
PLAYWRIGHT_CLI_SESSION="$(basename "$PWD")" claude

Run playwright-cli list to see every open session, and playwright-cli show to open a dashboard with a live view of each one. For the full worktree setup, including port blocks, see running many agents at once.

How to prove the browser check means something

Section titled “How to prove the browser check means something”

A browser check makes the agent’s claim observable. It proves the change only when the check could have failed and is kept where the agent cannot quietly weaken it.

CheckWhat it provesWho owns it
Red run before the fix, with the failure messageThe test detects this bugAuthor (agent); the reviewer reads the message
Green runs with --repeat-each=5 --retries=0The fix works and the test is not flakyAuthor (agent); CI reruns it
Test file unchanged between the red and green commitsThe agent fixed the code, not the testCI diff check; see protecting the oracle
E2E suite in CI with --fail-on-flaky-tests or zero retriesThe behaviour stays fixed on every later changeCI, blocking
Trace and screenshot attached to the pull requestA reviewer can replay what the browser did without rerunning itAuthor (agent)

The tech lead approves on that evidence. For the reviewing side of the loop, see reviewing an agent’s pull request.

What Playwright MCP and Playwright CLI cost in context

Section titled “What Playwright MCP and Playwright CLI cost in context”

Measured on 2026-09-26 by listing the tools of @playwright/mcp 0.0.82 over stdio:

ConfigurationToolsTool-schema JSON
Default25about 20,100 characters
--caps testing,network34about 25,700 characters
All seven capability groups72about 45,400 characters
Playwright CLI skill0 (shell commands)the SKILL.md body is 15,179 characters, loaded when the skill triggers

The schemas are the smaller cost. Claude Code 2.1.283 and codex-cli 0.157.1 both load MCP tool schemas on demand through tool search. The larger cost is the accessibility snapshot that Playwright MCP returns after every action: it grows with the size of the page and repeats with every click, so a ten-step flow on a content-heavy page pays for ten snapshots.

Ways to cut it, cheapest first:

  • Use Playwright CLI, which writes snapshots to files.
  • On Playwright MCP, ask for browser_snapshot with depth or a target element, or with filename to save it instead of returning it.
  • Start the server with --snapshot-mode none when the agent drives a known flow and does not need the page after each step.
  • Add only the --caps groups the task needs.
  • Push long browser sessions into a subagent so only its summary returns to the main session.

To measure it yourself in Claude Code, run /context before and after the pricing-page check, once with each driver. For the general pattern, see reducing MCP token cost.

Section titled “How popular are Playwright MCP and Playwright CLI?”

As of 2026-09-26: microsoft/playwright-mcp has 37.6k GitHub stars and is listed in the Official MCP Registry as io.github.microsoft/playwright-mcp 0.0.82; the playwright plugin has 319,887 installs on the Claude plugin directory (claude.com/plugins). microsoft/playwright-cli has 13.6k GitHub stars, and its skill shows 165,672 all-time installs in a third-party scrape of skills.sh (LinklyAI/best-skills, dated 2026-09-26; secondary, because skills.sh itself could not be read).

When Playwright MCP or Playwright CLI breaks

Section titled “When Playwright MCP or Playwright CLI breaks”

The browser fails to launch or reports a locked profile. Another agent or an earlier session holds the persistent profile. Recovery: add --isolated to the MCP entry, or give each agent its own --user-data-dir. For a stuck CLI session, run playwright-cli list, then playwright-cli close-all, and playwright-cli kill-all if processes remain.

No browser is installed. A fresh machine or CI container has no browser binaries, and the first navigation fails. Recovery: run npx playwright install chromium (add --with-deps on Linux CI images), or playwright-cli install-browser for the CLI. Choose a channel explicitly with --browser chrome|firefox|webkit|msedge.

The test passed against the wrong server. The dev server moved to the next free port, and the agent or reuseExistingServer hit another worktree’s server. Recovery: pin a port per worktree, derive baseURL and webServer.url from it, and ask the agent to prove which server it tested, as the worktree prompt does.

The agent edits the test until it passes. The red run was real; the green run came from a weaker assertion. Recovery: make “test file unchanged” a CI check, and put tests/e2e/ under a deny rule or CODEOWNERS as described in protecting the oracle.

The test is flaky. It passes locally and fails in CI, often on a fixed waitForTimeout or a CSS selector. Recovery: rerun with --repeat-each=10 --retries=0 and --trace=on, open the trace, and ask the agent to replace sleeps with web-first assertions and CSS selectors with role-based locators.

The login state is gone. --isolated discards cookies when the browser closes. Recovery: save a storage state once with Playwright’s setup project or playwright-cli state-save auth.json. On Playwright MCP, pass it with --storage-state auth.json; on Playwright CLI, run playwright-cli state-load auth.json. Keep that file out of Git; it holds session cookies.

The context window fills after a few clicks. Each MCP action returned a full snapshot of a large page. Recovery: switch that task to Playwright CLI, or use the snapshot limits in the previous section.

Where to go next with browser testing for agents

Section titled “Where to go next with browser testing for agents”