APIs, integration & security — in depth

Citation and Source Provenance in Web-Grounded Agent Responses

Web-grounded agents cite sources they don't actually verify.

Staff Writer · · 11 min read
Cover illustration for “Citation and Source Provenance in Web-Grounded Agent Responses”
Live Web Reasoning · October 1, 2026 · 11 min read · 2,383 words

Web-grounded agents are built on a simple promise: pull in live content from the internet, and every answer becomes checkable against a real source. That promise breaks down in a specific, measurable way. Retrieving a document and citing it correctly are two separate acts, and a model can complete the first without completing the second. It can pull a real page, generate a claim, attach a working URL, and still get the citation wrong, because the page is topically on-point but does not actually say what the claim says it says. A citation study from PricewaterhouseCoopers puts a number on this gap: the strongest frontier models keep link validity high and topical relevance solid, yet reach only 39 to 77% factual accuracy. The citation can look right while the claim behind it fails to hold up.

That range, 39 to 77%, is the whole argument in two numbers. A working link and a relevant source are necessary conditions for trustworthy attribution, not sufficient ones. An agent can pass every surface-level check, an accessible URL, a page about the right subject, and still fail the check that matters: whether the source says the specific thing the model claims it says. Fixing that gap means understanding where it opens up, where the failure modes get specific.

The specific ways citation breaks down in practice

Diagram: Five Citation Failure Types: From Fabrication to Staleness. Visualizes: Rank or sequence the five distinct citation failure modes identified in the article, ordered from most to least prevalent where data exists: (1) total fabrication —…

Citation failures sort into a small number of recognizable types, each with its own mechanism.

The largest category is outright invention: a URL that was never retrieved at all, fabricated by the model wholesale. A 2026 taxonomy of fabricated citations by Ansari found total fabrication responsible for 66% of cases, the dominant failure, followed by partial attribute corruption, identifier hijacking, placeholder hallucination, and semantic hallucination.

A second type is subtler and harder to catch by inspection: a source that is genuinely about the right subject but does not contain the specific fact being asserted. Research under the name PaperTrail, presented at CHI 2026 by Martin-Boyle and colleagues, isolates this failure by breaking both the model's answer and the source document into discrete claims and evidence units, then mapping which assertions are actually supported, which are unsupported, and which parts of the source were left out entirely. A citation can pass a human's quick glance, right topic, right domain, and still fail this decomposition.

A third type involves selective attribution rather than fabrication. Studies of conversational LLM agents show that claims backed by citations score highest on factuality, while ungrounded claims, the ones that likely come from the model's training data rather than retrieved content, score lowest. A meaningful share of claims trace back to URLs the agent actually retrieved during search but never cited at all. That pattern says something important: attribution in these systems is applied selectively, not enforced as a rule across every claim the model makes.

A fourth type involves content that shapes answers before it ever reaches citation. Untrusted content, pulled from webpages, documents, or repositories, can enter an agent's memory and persist across sessions, quietly shaping future answers without ever appearing as a cited source. A 2026 survey on evidence tracing and execution provenance by Wang and colleagues treats this as its own category, provenance-bearing evidence, precisely because the danger here is not a bad citation but a corrupted input that never gets citation scrutiny in the first place. Provenance, in other words, has to cover what gets into the corpus, not only what gets footnoted on the way out.

A fifth type is temporal rather than structural: a citation that was accurate the moment it was generated can become false later, because the page behind it changed. That failure looks nothing like hallucination on the surface, but it does the same damage to a reader trying to verify a claim.

What provenance means in an agent execution, beyond the inline citation

An inline citation is the visible tip of something that needs to run much deeper. Real provenance is a typed graph, one that connects a claim back through the specific tool call that produced it, the content that call retrieved, the URL that content came from, and the authority that justified using it in the first place. A citation with a working link satisfies almost none of that chain on its own.

Three properties make up that chain, and they need to be tracked separately rather than folded into one generic idea of "sourcing."

  • Decision provenance: which authority and which accepted facts actually governed an action, as distinct from evidence that was merely retrieved or sat nearby in the search results.
  • Execution provenance: every obligation tied to completion evidence that can be checked independently, so the agent's own claim of "done" is not treated as proof.
  • Change provenance: a record of which decisions and outputs depended on a source that has since been superseded, so recovery can invalidate exactly the stale work and nothing else.

A 2026 paper by Salas, "Correct Is Not Governed: Provenance Integrity in Agentic Workflows," makes this distinction concrete with a sharp example. An agent can land on the right business outcome while relying on an authority that had already expired or was still in draft form, can announce a task complete with no checkable evidence behind that claim, and can leave no trail showing which earlier work became stale once a policy changed. Correctness and governance are different measurements, and an agent can score well on the first while failing the second completely.

Tracing itself can happen at different levels of granularity, run-level, step-level, tool-call-level, parameter-level, claim-level, or down to individual tokens and spans, and the level chosen determines how precisely a failure can later be diagnosed and fixed. Multiple 2026 survey papers point to the W3C PROV-DM data model as the reference standard for representing these typed links between evidence and execution units. The specific standard matters less than the principle behind it: provenance needs a structure, not just a habit of attaching links.

The data ingestion layer shapes provenance before any citation is written

Provenance decisions get made long before an agent writes a sentence. Every RAG pipeline depends on live web content that postdates whatever the underlying model was trained on, and most conventional scraping tools fail on JavaScript-heavy pages or sites with anti-bot protections, returning blocked or partial content that cannot be traced back to any valid state of the source page. When that happens, the pipeline treats the URL as successfully retrieved even though the content indexed under that URL does not represent what the page actually says. That is a silent failure. Nothing in the pipeline flags it, and by the time a model cites that URL, the underlying error is invisible.

Purpose-built AI scraping tools exist specifically to close that gap, converting a given URL into clean, LLM-ready Markdown while keeping the source URL intact as the authoritative reference, removing the rendering failures and bot-detection blocks that corrupt content before it ever reaches an index. For structured extraction, schema-driven approaches, where a developer defines a JSON schema and the extraction layer returns data keyed to that schema, go a step further and attach provenance at the level of individual fields rather than whole documents. Some tools built for this kind of tracking return citations and bounding boxes tied to specific fields.

Corpus composition decisions carry the same weight. Which URLs get crawled, how content gets chunked and embedded, and which snapshot a retrieval query is scoped to, all determine which provenance chains are even possible once generation happens. A 2026 paper describing a self-evolving customer-support agent built at LinkedIn addresses this with a snapshot pointer: retrieval is scoped to the most recently ingested snapshot, but prior content is never deleted, so the system can always trace back which version of a source supported a given response. That kind of design decision, made entirely at the ingestion layer, is what makes later provenance claims possible or impossible. No amount of careful prompting at generation time can recover a chain that ingestion never preserved.

What architecture compensates for agentic tool-use breaking provenance

Retrieval used to be a fixed step that ran once, before generation started. In modern agentic architectures, retrieval is a tool the model decides to call, mid-reasoning, as many times as it judges necessary. That flexibility is valuable, but it multiplies the number of events that need a provenance record. Each tool call is a new event, and the record has to capture not just what came back, but why the agent chose to call retrieval at that point, with what parameters, and what the result actually was.

Wang and colleagues identify tool-use provenance as its own distinct research problem. Tool calls introduce state from outside the model, state the agent's reasoning depends on but cannot verify on its own, so the provenance layer has to independently log the call itself, the response it returned, and the chain connecting that response to whatever claim depends on it downstream. Trusting the model's own summary of what a tool returned defeats the purpose of tracking provenance.

Some recent research addresses this directly. ToolGate, presented at ACL 2026 by Liu and colleagues, introduces contract-grounded, verified tool execution: tool calls must satisfy preconditions declared in advance and return receipts that can be independently verified, rather than the pipeline simply accepting the model's account of what happened. Related work on auditable agents, including work by Nian and colleagues under that name and a paper by Rosen and Rosen on moving from agent loops to deterministic graphs, treats execution lineage and reproducibility as properties the architecture must guarantee from the start, not features bolted on for debugging later.

The Matrix framework described in the Salas paper is the clearest demonstration of what this buys in practice. Matrix externalizes provenance into a deterministic layer that sits around the probabilistic agent. The model keeps doing what models are good at, interpreting documents and proposing plans, while the framework owns stable identity, direct references, transition validity, materialization, declared preconditions, receipt matching, dependency traversal, bounded retry counters, and replay. In controlled comparisons, that architecture cut unnecessary task re-execution after a source changed from 18 tasks down to 3 across test seeds, achieved purely by tracking dependencies correctly instead of re-running everything out of caution, because the system could tell precisely which work actually depended on the changed source and which did not.

Source freshness as a provenance integrity condition, not a performance optimization

A citation that was accurate at the moment of retrieval can become false the moment the underlying page changes, and nothing about the citation itself will signal that anything is wrong. That makes tracking change a requirement for keeping the provenance chain intact.

RAG architecture has been shifting away from batch re-indexing, refreshing content on a nightly or hourly schedule, toward streaming re-indexing, where a document gets re-embedded as soon as it changes. Batch re-indexing is now treated as a design flaw rather than an acceptable tradeoff, because during the gap between refresh cycles the index can drift away from the live state of the source, with no record anywhere that the drift happened. A stale citation belongs in the same failure category as a hallucinated one: the agent cites a URL that no longer says what the citation implies, and there is no way for anyone reading the output to tell the difference without going and checking the live page themselves.

A newer set of tools treats webpage monitoring as something a pipeline can consume directly, rather than something a human has to notice and triage. Changes get represented as structured events, complete with hashes, diffs, and signed webhooks, and tools such as FluxProof offer this kind of evidence-bearing monitoring through standard REST and MCP endpoints. Some monitors go further and use an LLM to judge whether a detected change is semantically meaningful, a pricing update, policy change, or new announcement, versus noise such as ads, timestamps, and session IDs, notifying only when the change affects the content a citation depends on.

The change provenance property from the Salas paper formalizes this: a system needs a record of which decisions and outputs depend on a source that has since changed, so that recovery invalidates only the work that actually depends on it. Streaming re-indexing paired with structured change events is what makes that kind of selective recovery possible, provided the pipeline was built to track those dependencies in the first place.

A provenance-complete retrieval pipeline end to end

Provenance completeness is not a feature added at the point where a citation gets rendered on screen. It is a series of decisions made at every stage of the pipeline, and any one of those stages can break the chain if it is treated as an afterthought.

Ingestion has to produce, for every URL crawled, a content artifact stamped with its source URL, the timestamp it was retrieved, and the snapshot or version it belongs to. Partial or failed fetches need to be flagged explicitly rather than silently absorbed into the index as if they succeeded.

Chunking and embedding have to preserve enough metadata, source URL, page section, retrieval timestamp, that any chunk pulled during retrieval can be traced back, without ambiguity, to the exact page and position it came from.

Generation needs a mechanism that ties every claim the model produces to the specific chunk or chunks that support it, so that a claim without supporting content is visible as such rather than blending in with claims that are properly grounded.

Citation needs to be systematic rather than selective, covering every claim drawn from retrieved content, including the material that gets used in reasoning but currently tends to go uncited.

Change response needs a live connection back to the ingestion layer, so that when a source updates, the system knows precisely which prior claims and which prior work depended on the old version, and can invalidate that work without discarding everything else built on unaffected sources.

None of these stages compensates for failure at another. A generation step that ties claims to chunks does no good if the underlying chunk came from a malformed scrape that never represented the source page accurately in the first place. Provenance, built this way, is a property of the whole chain, tested at its weakest link rather than its most visible one.

Sources

  1. Correct Is Not Governed: Provenance Integrity in Agentic Workflows
  2. (PDF) PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
  3. From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
  4. Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
  5. Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
  6. Self-evolving Agentic Customer Support System at LinkedIn
  7. BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation

More in Live Web Reasoning