Crawl-and-Index API Integration With Pinecone and LlamaIndex
Separate fetch, orchestration, and storage layers for better RAG retrieval from live web data.

A working retrieval-augmented generation (RAG) system grounded in live web data needs three separate jobs done well: fetching clean web content, orchestrating chunking and embedding, and storing vectors for fast retrieval. You wire a crawl API, LlamaIndex, and Pinecone into one pipeline, and each tool does exactly one job. Get the fetch layer wrong and nothing downstream recovers, no matter how good the model answering the question is.
Why a crawl-and-index pipeline has three distinct jobs
You ground a RAG system in live web data when you run three jobs in sequence. Clean web content has to get fetched first. Then it has to be chunked and embedded in the right order. Finally the resulting vectors need a place to live where they can be searched quickly. Pipelines tend to break at the seams between these jobs, usually because someone merged two of them into one tool without realizing what got lost in the process.
Each job has a different failure signature, so you need to keep them separate. A fetch problem appears in a chunk as garbled or missing text. An orchestration problem is visible as chunks that split in the wrong place or embeddings that never get updated when source content changes. A storage problem occurs as slow queries or vectors that don't match the metadata filters applied at search time. Mixing the jobs together makes every one of those problems harder to isolate, because the symptom and the cause end up buried in the same piece of code.
The fetch layer is the one most teams shortchange. Raw HTML is not what a text splitter or an embedding model wants to see: it's full of navigation menus, ad slots, cookie banners, and script tags that have nothing to do with the actual content of the page. Pages rendered client-side with JavaScript often return blank or partial text to anything that isn't a full browser. Every one of these problems turns into a noisy chunk, and a noisy chunk turns into a bad embedding, and a bad embedding turns into a retrieval result that doesn't answer the question. The fetch layer sets the ceiling for everything else in the pipeline.
What LlamaIndex and Pinecone each do in this stack
Knowing where LlamaIndex's job ends and Pinecone's job begins heads off the most common integration mistake in this kind of pipeline: treating Pinecone like it can orchestrate, or treating LlamaIndex like it can store.
LlamaIndex is the orchestration layer. It provides the building blocks that move a document from its raw form into something retrievable: the VectorStoreIndex, the SentenceSplitter, the StorageContext, and the query engine. The packages involved in an Azure-flavored build include llama-index, llama-index-embeddings-azure-openai, llama-index-llms-azure-openai, llama-index-vector-stores-pinecone, and llama-index-readers-file, each handling one part of that flow. LlamaIndex supports more than 160 data connectors along with multiple chunking strategies, hybrid search, re-ranking, and query routing, so the orchestration surface is wide even in a simple build. The framework centers a RAG build on four jobs: ingestion, indexing, retrieval, and query control. Its open-source core is released under a permissive license. Picking an index type isn't a one-time decision made and forgotten: vector retrieval, keyword retrieval, tree indexes, and knowledge-graph approaches each fit different kinds of queries, so the right place to start is figuring out what questions users will actually be asking.
Pinecone is the storage layer. It's a managed vector database. It handles index creation, scaling, replication, and failover without anyone having to tune sharding or manage a Kubernetes cluster. Pinecone's own documentation frames this as letting teams focus on the RAG pipeline itself. A serverless deployment option scales with query volume and vector count, so it suits workloads where traffic is uneven. Namespaces let vectors be partitioned logically, useful for multi-tenant setups or for keeping separate content sources apart inside one index. Pinecone's API is proprietary and specific to Pinecone, even though its OpenAPI spec is published publicly, so if you build a pipeline tightly around it, you need to rewrite vector operations to move to a different store later.
The two layers connect through LlamaIndex's PineconeVectorStore integration, provided by the llama-index-vector-stores-pinecone package. That integration lets the orchestration layer write embeddings to Pinecone and query them back, and you need no custom glue code to hold the two systems together. Neither LlamaIndex nor Pinecone fetches content from the web, though, and that gap is exactly where the next layer in this stack has to come from.
The fetch layer's effect on retrieval quality
Garbage in, garbage out is the dominant failure mode in RAG pipelines built on web data, and it lives entirely at the fetch layer, before a single token reaches the embedding model.
These failure modes appear at the fetch layer in three specific forms. Navigation chrome, ads, and boilerplate end up embedded directly in chunks, because a chunker has no way to tell page content apart from site furniture once it's all just text. JavaScript-rendered content returns empty for crawlers that don't execute scripts, so entire pages can go missing from an index without any error being thrown. And re-crawling at scale turns into its own infrastructure project when built in-house: proxy rotation, anti-bot handling, and rendering each need separate ongoing maintenance.
The structural fix is to convert raw HTML into clean Markdown before any of it reaches the chunker. Markdown headings give a splitter natural seams to cut along, tables become parseable instead of a soup of div tags, and boilerplate gets stripped at the source. A recursive Markdown splitter, such as LlamaIndex's MarkdownNodeParser, behaves far more predictably on clean Markdown than on HTML-derived text, because chunk boundaries line up with the actual structure of the content.
So you should route web content through a crawl API built specifically for this job, rather than scrape HTML and patch it up after the fact. Rather than building JavaScript rendering, proxy rotation, and boilerplate stripping in-house, teams can use a crawl API designed to output clean Markdown ready for a recursive splitter, which eliminates all three failure modes before chunking even starts. Context.dev fits this role as the fetch layer: its API converts any URL into LLM-ready Markdown, crawls entire sites rather than single pages, and extracts structured data through a JSON schema when that's what the use case calls for. The output is built to feed directly into LlamaIndex's ingestion pipeline with no preprocessing step in between.
Setting up the environment and dependencies before writing a line of pipeline code
Before any pipeline code runs, three separate accounts and credential sets need to be in place: one for the crawl API, one for Pinecone, and one for whichever embedding model provider the pipeline will use. A Pinecone account is the first of these, and the free tier is enough to get started, since all that's needed at this stage is an API key and an index name, both available from the Pinecone console. You should put that API key in an.env file as PINECONE_API_KEY alongside PINECONE_INDEX_NAME, and you should never commit it to source control.
You need to set a handful of tuning parameters correctly before ingestion starts, because they affect storage cost, retrieval latency, and answer quality all at once. The text-embedding-3-small model supports higher dimensions than most RAG workloads actually need, and a reduced setting lowers storage costs and speeds up similarity search with little quality loss for typical use cases. Chunk size defaults to 1,000 tokens per chunk in this pipeline, and a small overlap between adjacent chunks prevents retrieval from losing content that happens to fall right on a chunk boundary.
On the Python side, the core dependencies are llama-index, llama-index-vector-stores-pinecone, llama-index-readers-file, and pinecone itself. Whatever embedding and LLM provider you choose for the build, you need to add its corresponding LlamaIndex package alongside these four, and the integration pattern stays the same regardless of which provider that ends up being. If you get all three sets of credentials and these packages in place before writing pipeline code, nothing in the steps that follow depends on stopping to go set up an account mid-build.
Step 1 (Fetching clean web content via the crawl API)
Here you implement the fetch layer in a single REST call, and it returns Markdown the rest of the pipeline can consume without any cleanup pass. A call to Context.dev's scrape endpoint for a single page looks like this:
POST /v1/web/scrape
{"url": ", "formats": {"markdown": true}}
The response body comes back as cleaned page content, and navigation and boilerplate are already stripped out. For a full site rather than a single page, the site-map or crawl endpoint enumerates the URLs first, and each one gets fetched afterward. Keeping discovery and extraction as two separate, observable stages means that if a bad chunk turns up later in retrieval, it's traceable back to a specific source URL instead of buried somewhere inside a single opaque crawl job.
The Markdown that comes back maps directly onto LlamaIndex's Document object. Page content becomes the document body, and the source URL becomes metadata that travels with every chunk you split from it later:
from llama_index.core import Document
doc = Document(
text=markdown_content,
metadata={
"source_url": url,
"content_hash": content_hash,
"updated_at": timestamp,
},
)
Three metadata fields matter enough to store on every document at this stage: the canonical URL, a content hash, and an updatedAt timestamp. A matching content hash means a page hasn't changed and doesn't need re-embedding, so these are what make re-crawls idempotent, and they're what make metadata filtering possible later at query time. Context.dev handles JavaScript rendering, proxy rotation, and anti-bot logic behind the API call, so whoever's calling it never has to know or care whether a given page was server-rendered or built client-side in the browser.
Step 2 (Chunking and embedding with LlamaIndex's IngestionPipeline)
LlamaIndex's IngestionPipeline is built for exactly this step: define the flow once, document loader, text splitter, embedding model, and let it run. The pipeline supports batching, caching, deduplication, and incremental updates, though each of those needs explicit configuration, such as passing in a persistent cache object and attaching a docstore.
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import MarkdownNodeParser
from llama_index.embeddings.azure_openai import AzureOpenAIEmbedding
pipeline = IngestionPipeline(
transformations=[
MarkdownNodeParser(),
AzureOpenAIEmbedding(dimensions=512),
],
)
nodes = pipeline.run(documents=[doc])
Chunking decisions here affect retrieval quality directly, and they're not something to leave on autopilot. If you start with a chunk size around 1,000 tokens and a small overlap, you get something reasonable for web content, because Markdown headings give a recursive splitter natural places to break that line up with the content's own structure. MarkdownNodeParser fits better than LlamaIndex's default SentenceSplitter when the input is clean Markdown coming out of a crawl API, because it respects heading hierarchy.
The crawl API sits upstream of LlamaIndex's document loader in this stack. It produces the clean, structured input that LlamaIndex's IngestionPipeline, its chunking strategies, and its embedding orchestration are all designed to consume, rather than having to compensate for messy input after the fact. Each chunk coming out of this step should carry the source URL and heading path as metadata, so the retrieval layer can surface citations back to a reader, and so Pinecone's metadata filters can scope a query down to a specific source when the question calls for it.


