APIs, integration & security — in depth

MCP Server Design for Web Data Tools

MCP servers need strong schema design and infrastructure to keep agents from hallucinating data.

Senior Writer · · 10 min read
Cover illustration for “MCP Server Design for Web Data Tools”
Tool and API Design · October 4, 2026 · 10 min read · 2,255 words

MCP (Model Context Protocol) is the standard that lets an AI agent find a tool, understand what it does, and call it, with no developer writing custom wrapper code for each one. That discovery property is the whole reason it beat plain REST APIs as the way agents reach outside tools. A REST endpoint needs a person to write a wrapper function, document each parameter by hand, parse the response, and translate errors into something an agent can act on. That glue code is specific to one provider, and it breaks the moment the API changes shape.

MCP skips that step. The protocol itself carries the typing and the invocation rules, so the agent already speaks the language the tool exposes. Adoption backs this up: seven of the top ten agent frameworks support MCP natively, and the rest get there through community adapters. Both the Python and TypeScript SDKs are maintained and current, which is a strong signal that this isn't a niche standard but the default one.

An MCP server exposes a small set of things an agent can enumerate on the spot: Resources, which are read-only data; Tools, which are actions the agent can trigger; and Prompts, which are reusable templates with structured arguments attached. A fourth primitive, Sampling, let a server send a completion request back to the client, but it's deprecated as of the MCP specification dated July 28, 2026. Everything that follows in this piece, about capability tiers, schema design, and output formats, assumes this foundation and builds on top of it.

Why web data tools fail inside agent loops

Getting MCP right at the protocol level is the easy part at this point. The failures that actually cost time and money appear one layer down, in what the tool fetches, what it hands back, and what it tells the agent about what just happened. The most counterintuitive failure looks like success: the tool call returns a success status on a page that's actually a CAPTCHA screen, and the agent treats that challenge page as real content, then tries to pull product data out of it. Nothing in the protocol flagged the problem, so the agent has no reason to doubt what it got.

When an agent falls apart on a live website, the planner is usually fine. The reasoning chain holds together. From the agent's point of view, a tool call that returns garbage looks identical to one that returns real data, unless the tool is built to tell the difference.

Other failure patterns follow the same shape. A model asked to extract data from a page it can't parse will sometimes invent a CSS selector that happens to look plausible. None of these are reasoning failures. They're infrastructure failures wearing a reasoning costume.

The quieter failure mode is that unchecked tool output wastes tokens. Every one of these failures traces back to the same root cause: the tool returned something, so the agent assumed the call worked. The protocol gave it no way to tell a clean response from a broken one, and a well-built web data tool is supposed to close that gap. A web scraping API designed for agent pipelines, like Context, handles bot detection, JavaScript rendering, and response validation at the infrastructure layer, so what reaches the agent is content that's already been checked, not a page that merely returned a status code.

The four capability tiers a web data MCP server can occupy

Web scraping MCP servers split into four capability tiers, and a server built for one tier can't stand in for another without real rework. The job at hand decides which tier is correct.

The first tier is static fetch: an HTTP request turned into plain text or Markdown. It's a good fit for server-rendered blogs and documentation pages, and it breaks the instant content depends on JavaScript to appear, or the response that comes back is a bot challenge instead of the real page.

The second tier is JavaScript rendering, where the page loads inside a browser or a rendering service before anything gets extracted. This brings client-rendered tables and product grids into the DOM where a tool can actually read them. Rendering the page doesn't guarantee a protected site hands over the real content rather than a block page, so this tier solves one problem without solving the other.

The third tier is full browser control: the agent can navigate between pages, click, type, wait for elements, switch tabs, and hold onto session state across steps. It still doesn't include a managed layer for getting past anti-bot systems. Full browser control gets you around a site; it doesn't get you past the gate in front of it.

The fourth tier is structured extraction, where the server hands back named fields or clean Markdown instead of a full HTML document.

There's also a runtime decision that applies across all four tiers: local versus hosted. Local MCP servers talk over stdio and run on the user's own machine. Hosted servers use Streamable HTTP transport and run the browser, the proxy, and the extraction layer somewhere remote. That choice decides who's on the hook for keeping the browser updated, managing IP reputation, handling anti-bot defenses, and keeping request logs. Protected sites running behind systems like Cloudflare, Akamai, or DataDome need proxy infrastructure, fingerprint handling, and challenge-solving no matter how smart the agent reasoning on top of it is. That's a fixed infrastructure cost, and no amount of language-model cleverness gets around it.

The design of these tiers, what they expose and what they quietly handle underneath, decides whether an MCP server acts as a thin pass-through or an actual data abstraction. A well-built tier can take on friction like anti-bot handling, session management, and response checking so the agent never has to reason through any of it. Platforms built as APIs for agent pipelines bake those layers in from the start, which keeps the MCP surface itself simple.

How tool schema design enables agents to call web data tools without human help

The schema a tool exposes is the contract the agent reasons over when it decides what to call and what arguments to pass. Get the schema wrong and the agent starts guessing, which produces made-up arguments and failures nobody notices until the output is already wrong.

Agents pass arguments by reading a schema and reasoning about it. A loosely typed endpoint that just takes a string in and hands a string back forces the agent to guess at parameter names and then parse whatever comes out the other side in freeform text. That guessing is exactly where hallucinated arguments and quiet failures come from.

Five properties separate a schema an agent can use on its own from one that needs a human standing by. An agent needs to enumerate what a tool can do and read typed parameter definitions at runtime, without a developer writing a custom wrapper for every single endpoint first. Every input and output needs an actual type attached to it. Output needs to come back as clean Markdown or structured JSON that keeps the meaning and drops the boilerplate, because the agent has to act on what it gets back, not just receive it and move on. When something fails, the tool needs to return a signal the agent can reason about and recover from, since an opaque failure just stalls the whole loop. And pricing needs to be predictable per call, because an agent running in a loop might call the same tool dozens of times in a single task, and a flat per-call price keeps that budget under control in a way that per-byte or per-proxy-gigabyte pricing never does.

Typed discovery and cost predictability affect the most in practice: typed discovery and cost predictability. Typed discovery is what makes the "no wrapper code" promise from the opening section actually real, rather than just a claim on a landing page. Cost predictability is what keeps an agent loop from quietly burning through a budget on retries against a blocked page, which turns out to be one of the single biggest cost drivers in production agent runs.

On top of those five properties, the minimal set of tools that covers almost everything an agent needs from the web is a three-way split: scrape for pulling one URL, crawl for working through a whole site, and map for discovering URLs without pulling their content. These three cover the vast majority of jobs, and most of what else exists is a specialized variant built on top of them. Output format, whether that's Markdown, HTML, JSON, or a screenshot, should be something the agent picks per call rather than something the server decides for it, so the agent can match the format to its task without a wasted second round-trip.

Schema-Driven Structured Extraction for Web Data

Clean Markdown removes page noise, but the agent still has to extract the fields it needs itself. If the agent's actual goal is five specific fields pulled from a hundred product pages, handing it clean page text still means it has to do the extraction work itself, one page at a time, burning tokens and getting slightly different results each time it tries.

Schema-driven extraction changes what the server is responsible for. Instead of returning a page, the server takes a field definition, something like "pull the name, price, rating, and availability from every product page", and hands back structured JSON that matches exactly that shape. When extraction is the actual job, a schema beats a block of raw text every time, because the schema tells the server precisely what counts as a correct answer.

This works because it puts the field definitions where the knowledge about them actually lives. The developer knows which fields matter for the product being built. The agent just needs to know how to call a typed tool correctly. The server handles the mechanics of finding those fields on the page and returning them in the right shape. Each piece does the part it's actually suited for instead of agents reverse-engineering structure out of a wall of text.

Structured extraction needs to be its own named tool inside the MCP server's schema. Agents find tools by reading their names and descriptions, not by exploring every possible combination of parameters a single tool might accept. Context.dev is one example of a platform built around this idea directly, offering schema-based extraction as a distinct capability rather than an option tucked inside a general scrape call. A JSON schema handed to the server means the server returns exactly the requested fields, and the agent can act on the result immediately.

Output Format and Content Cleaning Decisions in an Agent Loop

One noisy page is a rounding error. The same noisy output returned on every single call inside a multi-step agent loop turns into the dominant cost of the whole run, and it quietly degrades the quality of every reasoning step that follows.

Raw HTML is mostly navigation menus, cookie banners, and tracking scripts, so a model reading raw HTML spends real tokens wading through clutter instead of the content it was sent to find. Clean Markdown beats raw HTML at every stage of a language model pipeline, because it cuts token cost directly and strips out navigation, ads, and boilerplate before the content ever reaches the model. The same logic holds at the RAG ingestion step just as much as it holds for a single tool call.

A "main content only" setting, one that filters a page down to its actual semantic body, cuts more noise than almost any other single parameter without losing anything the agent actually needs. For those, screenshot output alongside Markdown covers what text can't capture on its own.

None of this is a rendering preference tucked away in settings. That cost occurs on every call, not just once. That cost compounds on every single call across the run, so a server's default output format matters just as much as which capability tier it sits in.

Designing MCP tools for live-web RAG so agents retrieve current data at query time

The most direct way to give an agent real-time web access is an API that fetches and cleans a page at the moment the agent asks, rather than one that serves answers out of an index built days or weeks earlier. A live-web RAG pipeline eliminates the freshness ceiling that static vector indexes run into, but only when the underlying tool schema actually supports the full live pipeline, not just a single fetch call.

Classic RAG goes stale because it answers out of a snapshot. A live-web RAG pipeline runs through seven stages: understanding the query, finding live sources, fetching and cleaning the content, chunking and embedding it, retrieving the top matches, grounding the generated answer in citations, and caching the result with a freshness window attached. An MCP tool needs to support at least the fetch-and-clean stage and the caching stage to be genuinely useful inside this kind of architecture.

Covering this well takes more than one capability. And some kind of monitor or cache with a freshness window matters too, since most production systems end up as a hybrid: fetching live when a query is time-sensitive, and reusing cached content when freshness isn't critical to the answer.

Context.dev covers this with a single API spanning scrape, crawl, batch, monitors, and web answers. Between the two, both pulling fresh content and keeping a cache current are covered without reaching for a second vendor to fill the gap.

Sources

  1. Context.dev: Web Scraping API for AI Agents & LLMs
  2. Web Scraping API Pricing: Plans from $0 | Context.dev

More in Tool and API Design