WebMCP: Making Websites Usable by AI Agents
WebMCP exposes website functions as callable tools for AI agents. Where Chrome, ChatGPT, Microsoft and Apple stand, and what it means for enterprises.
WebMCP exposes website functions as callable tools for AI agents. Where Chrome, ChatGPT, Microsoft and Apple stand, and what it means for enterprises.
Since 2025, browser vendors, model providers and standards bodies have been negotiating how a website should talk to an AI agent. WebMCP is currently the most visible answer to that question, and it is further along than most companies assume: Chrome runs an origin trial, ChatGPT uses it inside the browser of its desktop app, and Microsoft is a co-author of the specification. At the same time, Apple has filed an explicit opposition at WebKit. This page sets out what WebMCP actually is, who stands behind it, where the risks sit, and what a sensible first step looks like.
WebMCP is a proposed web standard that lets a website hand its own functions to an AI agent as named, typed and described tools. Instead of operating the page the way a human does — reading screenshots, scraping the DOM, guessing what a button means, computing click coordinates — the agent receives a list of callable functions with explicit input and output schemas: search_flights, add_to_cart, book_appointment, filter_results.
The difference is the same as between a remote control aimed at screenshots and an API. Agent actuation as practised today is non-deterministic and expensive: a shifted layout, a late-loading ad slot or a renamed label can break the whole automation loop, and every image analysis costs latency and tokens. One early adopter who wired WebMCP tools into browser test automation reported roughly a 90 percent reduction in token usage compared with the usual screenshot-action-screenshot cycle, alongside better repeatability.
One point matters for the framing: WebMCP complements the existing website, it does not replace it. Tools are registered in addition to the normal interface, they execute visibly on the page, and browsers without support continue to render the site unchanged. It is a progressive enhancement you can adopt incrementally, not a second front end to maintain.
The specification defines two ways to expose tools.
The declarative route annotates existing HTML forms with additional attributes. A search form gets a tool name and a description, and the agent then knows what the form is for and how to fill it in. For sites whose core functions already run through forms, this is the cheapest entry point.
The imperative route registers tools in JavaScript. Each tool carries a name, a description, an input schema expressed as JSON Schema, and an execute handler that calls the application logic you already have:
if (typeof document.modelContext?.registerTool === "function") {
await document.modelContext.registerTool({
name: "check_availability",
description: "Checks open appointment slots for a branch in a date range.",
inputSchema: {
type: "object",
properties: {
branch: { type: "string" },
from: { type: "string" },
to: { type: "string" },
},
required: ["branch", "from"],
},
annotations: { readOnlyHint: true },
execute: async ({ branch, from, to }) => findSlots(branch, from, to),
});
}
Three properties of the specification matter when you assess it. First, discovery: there is one standard way for a page to register its tools with an agent. Second, schemas: inputs and outputs are explicitly typed, which markedly reduces misinterpretation by the model. Third, state: the agent knows the current page context and therefore what it can act on right now.
The browser adds guard rails on top. Tools are only available in origin-isolated documents, so the document's origin stays stable for the lifetime of a tool. Both APIs are additionally gated by a permissions policy that defaults to the site's own origin, which means cross-origin embedded content cannot register tools unasked. Descriptions and outputs are subject to tight character budgets — among them 500 characters per tool description and 1,500 characters per tool output — which forces precise, compact definitions.
The Model Context Protocol, MCP, connects an AI application to a local or remote server. Those tools work regardless of whether a page is open: an MCP server can query a CRM or write records with no browser involved.
WebMCP, by contrast, operates entirely client-side, in the open document, inside the user's signed-in session. Nobody has to install a server or configure a connection first; the agent discovers the tools when it visits the page. That is the right mechanism whenever the human and the agent need to look at the same thing — editing a document, configuring a product, exploring a dashboard.
The two do not compete. A company can run an MCP server for session-independent integrations and additionally equip its website with WebMCP tools. Deciding what belongs where is an architecture question, not a front-end backlog item.
The picture is genuinely mixed, which is exactly why false assumptions circulate. Sorted by actor:
Google drives the specification and shipped first. WebMCP has been available in an origin trial since Chrome 149, and for local development the API can be switched on with a flag. The work was presented around the company's spring 2026 developer conference, alongside further building blocks for agentic browsing.
Microsoft is a co-author of the specification; standards bodies record the proposal as coming from Google and Microsoft. That is consistent with the line Microsoft has taken since introducing a Copilot mode in its own browser and agent features in Windows.
OpenAI has implemented WebMCP under the name "site tools" in the browser of the ChatGPT desktop application. Where an agent finds tools on a page, it can use them in the live, signed-in session; every invocation passes a safety review first, and the feature can be switched off in settings. OpenAI has also run a hackathon format with several platform providers to accelerate adoption across websites.
Apple, contrary to a widespread assumption, is not on board. The WebKit team filed a formal opposition in mid-2026, citing concerns about security, privacy, API design, duplication of existing platform features, portability, internationalisation and the choice of standards venue. This is not a formality: without WebKit there will be no native support on iOS for the foreseeable future.
Mozilla tracks the proposal in its standards positions but has not published a final position.
Standardisation itself runs in a W3C community group, accompanied by a design review from the Technical Architecture Group. WebMCP is therefore a serious proposal with two large implementers, but not a ratified standard. Investing today means investing in a likely future, not a guaranteed one.
The confusion about Apple's stance has an understandable cause: in summer 2026 Apple shipped its own MCP server for Safari, initially in the Technology Preview. It targets developers and lets coding agents inspect a real Safari session, read console output and network requests, and check pages.
That is a different thing from WebMCP. Apple's server drives a browser from the outside, for development and debugging. WebMCP works the other way round: the website itself offers tools, for end users. So Apple supports MCP as a protocol while rejecting WebMCP as a web platform feature. Reading those two announcements together produces the false conclusion that the industry has converged.
In practice: plan for Chromium browsers and agent applications as your target environment, not for blanket availability on iPhones. And expect the security concerns WebKit raised to shape how the specification develops.
The strategic question is not whether one particular API survives standardisation. It is this: what happens to our digital channel when a growing share of interactions is carried out not by people, but by their agents?
Three effects are already foreseeable.
Discoverability shifts from pages to capabilities. An agent with a task to complete will favour providers whose functions it can call reliably over those where it has to guess its way through multi-step interfaces. Offering tools simply makes you the cheaper path for agentic workflows.
Conversion paths get shorter and far more measurable. Booking, scheduling, requesting a quote, reordering: these are precisely the paths an agent completes on a customer's behalf. Modelling them cleanly as tools gives you a solid instrumentation of your own funnel as a by-product.
Support volume moves. A substantial share of service requests concerns transactions that could be self-served but that people fail to complete. Tools that start the right transaction directly and populate fields correctly reduce that volume measurably.
Against that sit risks nobody should wave away: opening functions to agents moves part of the control over your own customer interface into someone else's application. That trade-off belongs on the executive agenda, not in a front-end decision.
The specification's authors themselves note that language models are susceptible to indirect prompt injection, and that exposing native site functions creates risk that has to be understood and managed. Four points belong in every assessment.
External content is untrusted. Data that enters a tool response from outside — user comments, supplier text — must be marked as such so the agent treats it with additional scrutiny. Conversely, purely read-only operations can be flagged so that harmless queries do not trigger a confirmation prompt every time.
Permissions must not hang off the tool. A tool is an additional call path into existing logic, not a second, laxer entrance. The application's authentication, authorisation and input validation have to apply unchanged.
Stale rules become an operational risk. If a refund workflow is callable but the refund policy behind it is out of date, the agent will execute the wrong action cleanly and quickly. That is precisely the difference from a human who hesitates. Permission gaps can also become visible that were previously masked because nobody exercised them systematically.
Approvals need a defined stop. Anything that moves money, deletes data or communicates externally deserves an explicit confirmation step. Browser implementations provide mechanisms for this, but nobody else can decide for you which actions require it.
We therefore recommend treating WebMCP tools like a public API: with a threat model, logging, tests and periodic review. For companies already building AI governance, this is not a special case but another application of the same rules.
WebMCP is not a cure-all, and the parties involved say so openly. The API is designed for local workflows with a human in the loop, not for server-side bulk automation. Very complex interfaces require real refactoring, because application state and interface state have to be cleanly separated. And discoverability remains limited for now: an agent only learns that a site offers tools once it visits, which leaves open the question of a directory or some other signal.
There is also the obvious concern that agents bypass the commercial surface. If tools complete the core task directly, the intermediate pages where products are promoted, compared and cross-sold fall away. That question deserves a deliberate answer before the first tools go live.
We recommend a tightly scoped four-to-six-week pilot rather than a programme, following the same logic as our other pilot projects.
First, pick a single workflow with clear value and low blast radius: an availability check, a filter or search function, an appointment booking. Read-only tools are the best beginning because they demonstrate value without widening the security surface.
Then model two to five tools on top of existing application logic, with narrow schemas and terse descriptions. Test the result not only functionally but with scenario-based evaluations: does an agent complete typical requests reliably, and what happens with ambiguous or malicious input?
The output is a decision paper with numbers: success rate per scenario, steps required, cost per transaction, risks found. On that basis you can decide seriously whether and how far to expand.
WebMCP is not an isolated front-end topic. It touches three areas we work on in every mandate: prioritising use cases, governance for production operation, and the question of who owns a new capability long term.
We therefore treat it like any other pilot candidate rather than a special case pulled forward because it is topical. If your most important workflows are stalling internally, a customer service or document agent is the more effective first step. If your digital channel is the core of the business, WebMCP belongs in the next planning round, with a small, well-secured start.
For a first read on your own position, our AI readiness check gives you a structured baseline in a few minutes. To discuss the question concretely for your platform, book a strategy call, or look at how we connect strategy and governance with workshops.
Do I have to rebuild my website? No. WebMCP is registered in addition to what exists. The current interface stays as it is, and browsers without support notice nothing.
Does it work on iPhone? Not natively for the foreseeable future, because Apple opposed the API at WebKit. Agent applications that ship their own browser are not necessarily affected.
Is it stable enough for production? The specification is still moving and details may change. For a scoped, read-only use case the effort is nevertheless modest and the learning substantial.
What does it cost to start? The effort depends almost entirely on how cleanly your application logic is separated from your interface. In a modern application, two to five tools are a matter of days; in a grown interface, the same scope can take several weeks.