Five ways to drive a browser

· Code and every raw transcript: github.com/umutc/agent-browser-benchmark

I ran the same Claude Code agent through Claude in Chrome, Safari's new MCP server, Playwright on Chrome and WebKit, and Orca's built-in browser. 670 runs later, every surface could do the job. What separated them was cost, and a handful of sharp edges.

Safari 27 shipped an MCP server, which gave me a fifth way to let a coding agent use a browser. I already had Claude in Chrome, Playwright's MCP server, and the browser built into Orca, the tool I run my agents in. I did not want to pick one by feel, so I built a benchmark: the same agent, the same model, the same prompts, one browser surface per run, and a server that checks every answer against what actually happened on the page.

What I use now

For agent work in a clean browser, such as tests, scraping and form automation, I default to Playwright's MCP server. It never failed, and the engine is a flag. When token cost matters and the page has no native dialogs or drags, Safari's MCP server is surprisingly good. Claude in Chrome is for the times I need my own signed-in session, and I budget twice the tokens for it.

  • Playwright · ChromePlaywright · WebKit default
  • Safari MCP when tokens matter
  • Claude in Chrome my signed-in session

Results

670 runs, September 29–30, 2026
Surface Passedof 131 runs Probesof 22 Timemedian / run Inputmedian / run Fixedbefore any tool Cost134 runs Failed ontask: passes / runs
Claude in Chrome 126/131 20/22 26 s 122k 15.7k $22.5
  • HttpOnly cookie 0/3
  • print media 1/3
Playwright · Chrome 131/131 22/22 13 s 57k 10.6k $10.9 none
Orca browser 126/131 20/22 13 s 33k 6.4k $14.9
  • confirm() 1/3
  • print media 0/3
Safari MCP 127/131 21/22 13 s 38k 8.5k $10.7
  • confirm() 0/3
  • kanban drag 4/5
Playwright · WebKit 131/131 22/22 13 s 57k 10.6k $11.6 none

Probes counts capability probes passed in at least two of three runs. Time and Input are medians per run; Fixed is the input a reply costs before any tool is used. Cost is the list-price equivalent Claude Code reports for each surface's 134 runs; I ran on a Max subscription, and the whole benchmark used about 4% of a week's allowance. Bars are drawn to the largest value in each column.

Input tokens per run
  • Claude in Chrome122k
  • Playwright · Chrome57.2k
  • Orca browser32.8k
  • Safari MCP37.7k
  • Playwright · WebKit56.9k

Median over probes, realistic tasks and bug hunts.

Time per realistic task
  • Claude in Chrome39 s
  • Playwright · Chrome20 s
  • Orca browser27 s
  • Safari MCP22 s
  • Playwright · WebKit21 s

Median seconds for the eight multi-step tasks.

What I learned

Every surface can do the job. The difference is cost.

No surface fell below 96%, and Playwright passed all 131 runs on both engines. The spread is in time and tokens: the most expensive surface needed twice the time and nearly four times the input tokens of the cheapest.

Every run, one mark
Every run, one mark One mark per run, 134 per surface. Every run passed except: Claude in Chrome, HttpOnly cookie 0 of 3 and print media 1 of 3; Orca browser, JS dialogs 1 of 3 and print media 0 of 3; Safari MCP, JS dialogs 0 of 3 and kanban drag 4 of 5. Playwright on Chrome and on WebKit passed all 131 scored runs. baseline 22 capability probes × 3 8 realistic tasks × 5 5 bug hunts × 5 passed Claude in Chrome HttpOnly cookie 0/3 print media 1/3 126/131 Playwright · Chrome 131/131 Orca browser JS dialogs 1/3 print media 0/3 126/131 Safari MCP JS dialogs 0/3 kanban drag 4/5 127/131 Playwright · WebKit 131/131

Pass counts per task from the run summary, grouped by tier; within a task, passes are drawn first. Baseline runs answer without tools and are not counted in the totals.

Every task, pass count per surface · 36 tasks, 4 tiers
Passes / runs. Highlighted cells dropped at least one run; tasks with a drop are listed first in each tier.
Task Claude in Chrome Playwright · Chrome Orca browser Safari MCP Playwright · WebKit
Capability probes, 3 runs each
emulate print media1/33/30/33/33/3
HttpOnly cookie0/33/33/33/33/3
JS dialogs (confirm + prompt)3/33/31/30/33/3
click3/33/33/33/33/3
console messages3/33/33/33/33/3
drag and drop (HTML5)3/33/33/33/33/3
file download3/33/33/33/33/3
file upload3/33/33/33/33/3
form controls3/33/33/33/33/3
hover3/33/33/33/33/3
iframe (cross origin)3/33/33/33/33/3
iframe (same origin)3/33/33/33/33/3
infinite scroll3/33/33/33/33/3
keyboard shortcut3/33/33/33/33/3
network request body3/33/33/33/33/3
new tab / popup3/33/33/33/33/3
read text3/33/33/33/33/3
resize viewport3/33/33/33/33/3
run JavaScript3/33/33/33/33/3
screenshot + vision (below the fold)3/33/33/33/33/3
shadow DOM3/33/33/33/33/3
type and submit3/33/33/33/33/3
Realistic tasks, 5 runs each
pointer-based drag and drop5/55/55/54/55/5
login + 2FA across tabs + settings5/55/55/55/55/5
multi-field form with validation5/55/55/55/55/5
multi-page data extraction5/55/55/55/55/5
multi-step wizard + custom date picker5/55/55/55/55/5
overlays, modals, toasts5/55/55/55/55/5
paginated table aggregation5/55/55/55/55/5
SPA filters + async results5/55/55/55/55/5
Bug hunts, 5 runs each
accessibility audit5/55/55/55/55/5
engine-specific breakage5/55/55/55/55/5
find console errors per action5/55/55/55/55/5
find failing API calls5/55/55/55/55/5
mobile layout bugs5/55/55/55/55/5
Baseline, no tools
baseline3/33/33/33/33/3
  • Claude in ChromeChrome 154, my own profile
    extension

    Claude in Chrome costs about twice as much

    A median run took 26 seconds and 122k input tokens, against 13 seconds and 33k to 57k for the other four. It reaches for screenshots more often (a median of 4.5 images per realistic task), its fixed context is the largest at 16k tokens, and each tool call round-trips through the extension. It also could not read an HttpOnly cookie in any run. None of this measures what it is actually for: driving my own browser, where I am already signed in. For that it is still the only option here.

  • Safari MCPSafari 27
    safaridriver --mcp

    Safari MCP is the leanest, with two gaps

    It was the cheapest surface overall and returned the most compact page reads, a median of 0.8 KB per probe. Two things are missing. A native confirm() left its click call hanging until the four-minute timeout, in all three dialog runs. And it has no press-move-release gesture, so on the kanban board the agent had to fake the drag by dispatching pointer events from page JavaScript. Only 70% of its passing interactions reached the page as real input (event.isTrusted), against 100% for every other surface.

  • Orca browserChromium 150 (Electron)
    orca CLI through Bash

    Orca's browser is fragile around the host app

    A native confirm() dropped the Orca runtime connection, so every command after it failed; one of three runs got through with a workaround. set media print reported success without changing the media type. Its speed also depended on the Orca window itself: its viewport follows the panel size, and once the harness pre-opened a pinned tab and the workspace was on screen, its median run time halved.

  • Playwright · Chrome Playwright · WebKit @playwright/mcp 0.0.83

    The interface matters more than the engine

    Playwright on Chrome and on WebKit finished within a second of each other on every tier. The bug hunt built around engine differences came out exactly as the engines dictate: both WebKit surfaces found three broken widgets, because two of them call Chromium-only APIs, and all three Chromium surfaces found one.

  • Claude in Chrome Orca browser no print-media emulation

    The agents said when they took a shortcut

    Chrome and Orca have no way to emulate print media, so the agents improvised: one rewrote the page's print CSS and fired beforeprint, the other rendered a PDF to trigger the print handler. Both reached the right code and explained the workaround in their notes. The checker still fails those runs, because the task asked for real print media. They are the only five cases where an agent claimed success and failed.

The setup

Every run is a fresh claude -p process with exactly one browser surface. --strict-mcp-config loads a single MCP server, the only built-in tool is Read, and --setting-sources project keeps my personal instructions out of the run. Without Bash or WebFetch, the agent cannot quietly curl the page instead of using the browser. Orca's browser is a command-line tool, so that lane gets Bash restricted to orca browser commands and nothing else.

The five surfaces
SurfaceEngineInterface
Claude in ChromeChrome 154, my own profileExtension, built into Claude Code
Playwright · ChromeChrome 154, fresh profile@playwright/mcp 0.0.83
Orca browserChromium 150 (Electron)orca CLI through Bash
Safari MCPSafari 27safaridriver --mcp
Playwright · WebKitWebKit 26.6 (Playwright build)@playwright/mcp 0.0.83

The pages come from a local fixture server. Each run gets its own origin (run-<id>.localhost), so cookies and storage start empty in every browser, even in my everyday Chrome profile. The expected answers are derived from a secret that only exists in the runner's memory, so an agent that reads the benchmark's source still cannot compute them. There are 36 tasks in four tiers:

  • 22× 3 runs

    22 capability probes, three runs each: one skill per page, from clicking a button to a cross-origin iframe, a confirm() dialog, an HttpOnly cookie, a code drawn on a canvas below the fold, and print media.

  • 8× 5 runs

    8 realistic tasks, five runs each: a checkout with validation, a filtered catalog, a paginated table, a wizard with a custom date picker, a login with a code from a second tab, detail-page scraping, a shop full of pop-ups, and a pointer-driven kanban board.

  • 5× 5 runs

    5 bug hunts, five runs each: pages with planted problems that move every run, from failing API calls behind normal-looking widgets to widgets that only break in WebKit.

  • 1× 3 runs

    A baseline that answers without tools, to measure what each surface costs before it does anything.

66 + 40 + 25 + 3 = 134 runs per surface · × 5 surfaces = 670

What went wrong while measuring

The benchmark found bugs in itself before it found anything about the browsers, and I think those are worth listing:

  • fixture

    Two fixture bugs. Before the full run, a scripted Playwright solution solved every task in both Chrome and WebKit. It caught a pop-up whose display: flex overrode the hidden attribute, which would have failed every agent, and a console page where two identical WebKit exceptions were reported as one, so the correct answer would have looked wrong.

  • viewport

    A resized window. On the “make the viewport phone-width” probe, Claude in Chrome shrank my real Chrome window to 500 pixels and never restored it. The next four Chrome runs started narrow. The viewport beacon caught it; I discarded those runs, and the harness now restores window bounds and refuses runs that start narrower than 1700 pixels.

  • runner

    A false rate limit. An API 529 Overloaded error looked like a usage limit to the first version of the runner, which then planned to sleep for four and a half hours.

  • safety filter

    A safety filter. The HttpOnly-cookie probe tripped the API's cyber safeguard in 12 of its 15 runs. “Read this session cookie” looks a lot like cookie theft. The refused turn was retracted and re-run on an older model, so that probe mixes models; the other 658 runs used Opus 5.5 only.

The first 174 runs ran one at a time. For the rest I ran all five surfaces on the same task at the same moment, which was about five times faster and arguably fairer, since every surface saw the same machine load and the same API conditions. For four of the five, times on matching tasks moved by at most 6%.

Limits

  • One model at one effort setting. The tasks turned out easy for it, so this benchmark separates cost and edge cases better than general competence. A smaller model would probably spread the pass rates.
  • Claude Code ships its own system-prompt guidance for Claude in Chrome; the others rely on their tool descriptions.
  • Engine builds differ, and Playwright's WebKit is not the engine inside Safari 27.
  • Local fixture pages only: no sign-in walls, bot detection or very heavy real-world pages.

Sources

  1. umutc/agent-browser-benchmark: the harness, fixture site, task definitions, all 670 run records and every raw agent transcript (MIT). The full report with per-task tables is results/main/report.html.
  2. Introducing the Safari MCP server for web developers, WebKit blog.
  3. @playwright/mcp on npm; this benchmark used version 0.0.83.
  4. Claude Code, the agent every run used, version 2.1.285.

I designed and ran this benchmark with Claude Code as a pair. It wrote most of the harness and the first draft of this post; I reviewed the design, the anomalies and the conclusions.