Five ways to drive a browser
I ran the same Claude Code agent through Claude in Chrome, Safari's new MCP server, Playwright on Chrome and WebKit, and Orca's built-in browser. 670 runs later, every surface could do the job. What separated them was cost, and a handful of sharp edges.
Safari 27 shipped an MCP server, which gave me a fifth way to let a coding agent use a browser. I already had Claude in Chrome, Playwright's MCP server, and the browser built into Orca, the tool I run my agents in. I did not want to pick one by feel, so I built a benchmark: the same agent, the same model, the same prompts, one browser surface per run, and a server that checks every answer against what actually happened on the page.
What I use now
For agent work in a clean browser, such as tests, scraping and form automation, I default to Playwright's MCP server. It never failed, and the engine is a flag. When token cost matters and the page has no native dialogs or drags, Safari's MCP server is surprisingly good. Claude in Chrome is for the times I need my own signed-in session, and I budget twice the tokens for it.
- Playwright · ChromePlaywright · WebKit default
- Safari MCP when tokens matter
- Claude in Chrome my signed-in session
Results
| Surface | Passedof 131 runs | Probesof 22 | Timemedian / run | Inputmedian / run | Fixedbefore any tool | Cost134 runs | Failed ontask: passes / runs |
|---|---|---|---|---|---|---|---|
| Claude in Chrome | 126/131 | 20/22 | 26 s | 122k | 15.7k | $22.5 |
|
| Playwright · Chrome | 131/131 | 22/22 | 13 s | 57k | 10.6k | $10.9 | none |
| Orca browser | 126/131 | 20/22 | 13 s | 33k | 6.4k | $14.9 |
|
| Safari MCP | 127/131 | 21/22 | 13 s | 38k | 8.5k | $10.7 |
|
| Playwright · WebKit | 131/131 | 22/22 | 13 s | 57k | 10.6k | $11.6 | none |
Probes counts capability probes passed in at least two of three runs. Time and Input are medians per run; Fixed is the input a reply costs before any tool is used. Cost is the list-price equivalent Claude Code reports for each surface's 134 runs; I ran on a Max subscription, and the whole benchmark used about 4% of a week's allowance. Bars are drawn to the largest value in each column.
Median over probes, realistic tasks and bug hunts.
Median seconds for the eight multi-step tasks.
What I learned
Every surface can do the job. The difference is cost.
No surface fell below 96%, and Playwright passed all 131 runs on both engines. The spread is in time and tokens: the most expensive surface needed twice the time and nearly four times the input tokens of the cheapest.
Pass counts per task from the run summary, grouped by tier; within a task, passes are drawn first. Baseline runs answer without tools and are not counted in the totals.
Every task, pass count per surface · 36 tasks, 4 tiers
| Task | Claude in Chrome | Playwright · Chrome | Orca browser | Safari MCP | Playwright · WebKit |
|---|---|---|---|---|---|
| Capability probes, 3 runs each | |||||
| emulate print media | 1/3 | 3/3 | 0/3 | 3/3 | 3/3 |
| HttpOnly cookie | 0/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| JS dialogs (confirm + prompt) | 3/3 | 3/3 | 1/3 | 0/3 | 3/3 |
| click | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| console messages | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| drag and drop (HTML5) | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| file download | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| file upload | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| form controls | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| hover | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| iframe (cross origin) | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| iframe (same origin) | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| infinite scroll | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| keyboard shortcut | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| network request body | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| new tab / popup | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| read text | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| resize viewport | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| run JavaScript | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| screenshot + vision (below the fold) | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| shadow DOM | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| type and submit | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Realistic tasks, 5 runs each | |||||
| pointer-based drag and drop | 5/5 | 5/5 | 5/5 | 4/5 | 5/5 |
| login + 2FA across tabs + settings | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| multi-field form with validation | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| multi-page data extraction | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| multi-step wizard + custom date picker | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| overlays, modals, toasts | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| paginated table aggregation | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| SPA filters + async results | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| Bug hunts, 5 runs each | |||||
| accessibility audit | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| engine-specific breakage | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| find console errors per action | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| find failing API calls | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| mobile layout bugs | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| Baseline, no tools | |||||
| baseline | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
-
Claude in ChromeChrome 154, my own profile
extensionClaude in Chrome costs about twice as much
A median run took 26 seconds and 122k input tokens, against 13 seconds and 33k to 57k for the other four. It reaches for screenshots more often (a median of 4.5 images per realistic task), its fixed context is the largest at 16k tokens, and each tool call round-trips through the extension. It also could not read an HttpOnly cookie in any run. None of this measures what it is actually for: driving my own browser, where I am already signed in. For that it is still the only option here.
-
Safari MCPSafari 27
safaridriver --mcpSafari MCP is the leanest, with two gaps
It was the cheapest surface overall and returned the most compact page reads, a median of 0.8 KB per probe. Two things are missing. A native
confirm()left its click call hanging until the four-minute timeout, in all three dialog runs. And it has no press-move-release gesture, so on the kanban board the agent had to fake the drag by dispatching pointer events from page JavaScript. Only 70% of its passing interactions reached the page as real input (event.isTrusted), against 100% for every other surface. -
Orca browserChromium 150 (Electron)
orca CLI through BashOrca's browser is fragile around the host app
A native
confirm()dropped the Orca runtime connection, so every command after it failed; one of three runs got through with a workaround.set media printreported success without changing the media type. Its speed also depended on the Orca window itself: its viewport follows the panel size, and once the harness pre-opened a pinned tab and the workspace was on screen, its median run time halved. -
Playwright · Chrome Playwright · WebKit @playwright/mcp 0.0.83
The interface matters more than the engine
Playwright on Chrome and on WebKit finished within a second of each other on every tier. The bug hunt built around engine differences came out exactly as the engines dictate: both WebKit surfaces found three broken widgets, because two of them call Chromium-only APIs, and all three Chromium surfaces found one.
-
Claude in Chrome Orca browser no print-media emulation
The agents said when they took a shortcut
Chrome and Orca have no way to emulate print media, so the agents improvised: one rewrote the page's print CSS and fired
beforeprint, the other rendered a PDF to trigger the print handler. Both reached the right code and explained the workaround in their notes. The checker still fails those runs, because the task asked for real print media. They are the only five cases where an agent claimed success and failed.
The setup
Every run is a fresh claude -p process with exactly one browser surface. --strict-mcp-config loads a single MCP server, the only built-in tool is Read, and --setting-sources project keeps my personal instructions out of the run. Without Bash or WebFetch, the agent cannot quietly curl the page instead of using the browser. Orca's browser is a command-line tool, so that lane gets Bash restricted to orca browser commands and nothing else.
| Surface | Engine | Interface |
|---|---|---|
| Claude in Chrome | Chrome 154, my own profile | Extension, built into Claude Code |
| Playwright · Chrome | Chrome 154, fresh profile | @playwright/mcp 0.0.83 |
| Orca browser | Chromium 150 (Electron) | orca CLI through Bash |
| Safari MCP | Safari 27 | safaridriver --mcp |
| Playwright · WebKit | WebKit 26.6 (Playwright build) | @playwright/mcp 0.0.83 |
The pages come from a local fixture server. Each run gets its own origin (run-<id>.localhost), so cookies and storage start empty in every browser, even in my everyday Chrome profile. The expected answers are derived from a secret that only exists in the runner's memory, so an agent that reads the benchmark's source still cannot compute them. There are 36 tasks in four tiers:
- 22× 3 runs
22 capability probes, three runs each: one skill per page, from clicking a button to a cross-origin iframe, a
confirm()dialog, an HttpOnly cookie, a code drawn on a canvas below the fold, and print media. - 8× 5 runs
8 realistic tasks, five runs each: a checkout with validation, a filtered catalog, a paginated table, a wizard with a custom date picker, a login with a code from a second tab, detail-page scraping, a shop full of pop-ups, and a pointer-driven kanban board.
- 5× 5 runs
5 bug hunts, five runs each: pages with planted problems that move every run, from failing API calls behind normal-looking widgets to widgets that only break in WebKit.
- 1× 3 runs
A baseline that answers without tools, to measure what each surface costs before it does anything.
66 + 40 + 25 + 3 = 134 runs per surface · × 5 surfaces = 670
What went wrong while measuring
The benchmark found bugs in itself before it found anything about the browsers, and I think those are worth listing:
- fixture
Two fixture bugs. Before the full run, a scripted Playwright solution solved every task in both Chrome and WebKit. It caught a pop-up whose
display: flexoverrode thehiddenattribute, which would have failed every agent, and a console page where two identical WebKit exceptions were reported as one, so the correct answer would have looked wrong. - viewport
A resized window. On the “make the viewport phone-width” probe, Claude in Chrome shrank my real Chrome window to 500 pixels and never restored it. The next four Chrome runs started narrow. The viewport beacon caught it; I discarded those runs, and the harness now restores window bounds and refuses runs that start narrower than 1700 pixels.
- runner
A false rate limit. An API
529 Overloadederror looked like a usage limit to the first version of the runner, which then planned to sleep for four and a half hours. - safety filter
A safety filter. The HttpOnly-cookie probe tripped the API's cyber safeguard in 12 of its 15 runs. “Read this session cookie” looks a lot like cookie theft. The refused turn was retracted and re-run on an older model, so that probe mixes models; the other 658 runs used Opus 5.5 only.
The first 174 runs ran one at a time. For the rest I ran all five surfaces on the same task at the same moment, which was about five times faster and arguably fairer, since every surface saw the same machine load and the same API conditions. For four of the five, times on matching tasks moved by at most 6%.
Limits
- One model at one effort setting. The tasks turned out easy for it, so this benchmark separates cost and edge cases better than general competence. A smaller model would probably spread the pass rates.
- Claude Code ships its own system-prompt guidance for Claude in Chrome; the others rely on their tool descriptions.
- Engine builds differ, and Playwright's WebKit is not the engine inside Safari 27.
- Local fixture pages only: no sign-in walls, bot detection or very heavy real-world pages.
Sources
- umutc/agent-browser-benchmark: the harness, fixture site, task definitions, all 670 run records and every raw agent transcript (MIT). The full report with per-task tables is
results/main/report.html. - Introducing the Safari MCP server for web developers, WebKit blog.
- @playwright/mcp on npm; this benchmark used version 0.0.83.
- Claude Code, the agent every run used, version 2.1.285.
I designed and ran this benchmark with Claude Code as a pair. It wrote most of the harness and the first draft of this post; I reviewed the design, the anomalies and the conclusions.