Infra

WebMCP: I Tested Chrome’s Tool API for Browser Agents

WebMCP: I Tested Chrome’s Tool API for Browser Agents
WebMCP lets a web page hand tools to a browser agent. I ran it in Chromium 153, hit one real bug, and sketched a half-day Trellis experiment.

WebMCP is a proposed browser API that lets a web page give an AI agent a list of tools it can call, so the agent runs search_my_listings instead of squinting at buttons and screenshots. It is an open spec from the W3C Web Machine Learning community group, with Google and Microsoft authors, and Chrome already has it behind two flags. It is not the MCP protocol you use in Claude Desktop or Cursor, even if the name sounds like it.

The news is that Google runs an origin trial from Chrome 149 to 156 and says it wants to ship in 157, while WebKit formally opposes the idea and Mozilla is neutral. In this post I go over what the API looks like, who is for it and who is against, the one benchmark I found, the security gap that worries me, and the half-day experiment I would run at Trellis. My verdict in one line: try it on one read-only thing this month, and keep your hands off everything that writes.

Here is why I care. A Trellis user lives in a React app full of tables, filters and buttons. Today a browser assistant that wants “my listings with fewer than five images” has to click around that UI like a tourist. With WebMCP our page could just offer a tool called search_my_listings. Reads, I can picture. Writes to listings, pricing or ad spend stay out of my plans until the spec has a real consent model, and it does not have one yet.

I also ran it. In Chromium 153.0.8010.12 with two feature flags on, I registered one tool, asked the page for its tool list and got the name back. Then I called the tool with a plain object, the way the README shows it, and the browser answered “UnknownError: Failed to parse input arguments”. Same call with JSON.stringify around the object worked. So the spec and the build disagree on the first example.

What is WebMCP, and what is it not?

A page calls document.modelContext.registerTool() with a name, a description, a JSON schema for the inputs and an execute function. The browser’s built-in agent can then list those tools and call them, instead of scraping the DOM or reading screenshots. It is document.modelContext, not navigator.modelContext, though blog posts use both. The explainer in the community group repo is the source for all of this.

There is a second form with no JavaScript: toolname and tooldescription on a

, toolparamdescription on each input. The declarative explainer says schema synthesis is still a TODO, and the README calls declarative “not part of the specification at this time”. Chrome’s docs document it anyway.

Now the name. The spec says it “derives direct inspiration from MCP and shares a common vocabulary” (tools, schemas, parameters), and in the same document it lists adopting MCP directly as a rejected alternative, because MCP has no notion of origins, permissions or tab lifecycle. Mozilla’s review is shorter: “There is no MCP here.” Google’s own comparison page says MCP is persistent and backend-side, while WebMCP tools die with the tab and talk only to the browser’s built-in agent. So if you already run an MCP server, this is not a replacement for it. I wrote about running one for Amazon ads in my Amazon PPC MCP server post, and nothing in there goes away.

What did I actually run?

Only this. Playwright’s Chromium 153.0.8010.12, headless, on example.com. Without flags, document.modelContext is undefined. With --enable-features=WebMCPTesting --enable-blink-features=WebMCP the object appears, and its prototype has exactly four members: registerTool, getTools, executeTool and ontoolchange.

The tool was the README’s add-todo, and getTools() returned ['add-todo']. The call with the JSON string came back as a JSON string too: {"content":[{"type":"text","text":"added milk"}]}.

That is the whole test. I did not drive a real agent, touch the origin trial token or the declarative path, or run any Trellis code. The plumbing answers. Whether an assistant picks the right tool, I cannot say.

Pin your Chrome version when you test.

Abstract editorial illustration of a browser window with three amber doors on its edge, one open, and teal lines linking to a small panel, flat vector style on navy.
A page offers a few doors to the agent. Which doors you open is the whole decision.

Who is shipping WebMCP, and who is not?

The dates below are Google’s own statements, so weigh them that way. Google’s Intent to Experiment on blink-dev says an origin trial runs from Chrome 149 to 156, with a shipping target of Chrome 157.

The other engines are where it gets interesting. WebKit’s standards position #670 is OPPOSE, posted June 3. Their argument: an agent acting for a user is assistive technology, so fix HTML and ARIA instead of building a parallel layer, and the spec has no consent or reversibility model. Mozilla’s position #1412 is neutral: they see the benefit, they worry that hostile sites will tar-pit agents, hide prompt injection or harvest data through tool inputs, and later comments say they care about the imperative API, not the declarative one. The W3C TAG’s design review #1238 was still open when a maintainer commented on October 6.

My read: WebKit’s argument is the better one on principle, and Chrome’s has code behind it today. Half a day is cheap, a roadmap line is not.

The one hint of real use I found is a comment in that TAG thread from Shopify’s Yoav Weiss, who says he is deploying origin-trial tools on a large number of storefronts. He gives no count, and I found none.

Is a tool call really cheaper than clicking?

Maybe. There is exactly one benchmark I found, WindTunnel from nekuda-ai, and it is a vendor benchmark where the vendor wrote the WebMCP tools by hand, which their own SPEC discloses as a conflict of interest. Its README table says Sonnet 5 solved 49 of 49 tasks through WebMCP at a median of $0.009, against 48 of 49 at $0.210 for DOM plus vision. I found no independent replication.

A function call with a schema is surely less work for a model than a 64,424-token screenshot-and-DOM loop, so the direction looks right. But the 100% is partly a product of the authors writing the tools. Treat the 3x to 47x claim as the vendor’s.

The security part Trellis cannot skip

Read the security and privacy questionnaire in the repo. The authors admit there is no normative guidance yet for sensitive tools. Tool output flows into the agent, so anything an attacker can get into a field you return, say a review text or a buyer message, is a prompt injection path. Google’s secure tools guide names that as the main threat and offers untrustedContentHint, consequentialHint and readOnlyHint. The spec issue for the consequential hint is still open, and readOnlyHint is advisory.

That is the reason my line sits where it sits. A wrong price or ad budget costs a seller real money in minutes. Until the browser itself asks the human, with a screen the page cannot style, no write tool of mine goes live. No toolautosubmit on any form, either.

The half-day experiment

Bon. Here is what I would do, and what I will actually do if I find the half day.

  1. Pick one existing read-only lookup, for us the listing search that the UI already calls.
  2. Feature-detect document.modelContext. If it is missing, load the MIT-licensed @mcp-b/webmcp-polyfill, which pulled 169,652 downloads last week by npm’s own API.
  3. Register one tool, search_my_listings, with readOnlyHint on, wrapping that same function. Or, if you have a plain search form, add the three attributes and nothing else.
  4. Test with Google’s inspector from the webmcp-tools repo (Apache-2.0, 17 demos), and log every call to the tool with a timestamp and the user agent.
  5. Leave it on two weeks and read the log.

Mettons the log is empty after two weeks. Then I learned, for half a day of work, that my sellers do not use built-in browser agents yet.

What I still do not know is the part that decides everything. Which agents really consume these tools in general availability, today, outside a demo? The repo’s implementation-status file lists a few, and I verified none. If you have a site with tools registered and a log of who called them, I want to see that number more than anything else in this post.

Dominic Plouffe