Build a local PDF Q&A pipeline with Ollama, LangChain, and Chroma that returns both an answer and the source chunks behind it. The citations make retrieval and interpretation errors visible instead of burying them inside a fluent response.

What Changed in LangChain v1

Older examples import RetrievalQA, LLMChain, text splitters, Ollama, and Chroma from the top-level langchain package. LangChain v1 reduced that namespace. Legacy chains moved to langchain-classic, while maintained integrations are installed in separate packages.

This guide avoids the legacy RetrievalQA abstraction and uses current integration packages directly:

python3 -m venv pdf-qa-env
source pdf-qa-env/bin/activate
python -m pip install --upgrade pip
python -m pip install \
  langchain-community \
  langchain-chroma \
  langchain-ollama \
  langchain-text-splitters \
  pypdf

Pull one model for answers and one for embeddings:

ollama pull llama3.1:8b
ollama pull nomic-embed-text

Pin package and model versions for a durable deployment. The commands above install whatever versions are current when run; that is convenient for a tutorial but not reproducible production configuration.

Extract the PDF and Preserve Page Metadata

PyPDFLoader extracts text from text-based PDFs. It does not make a scanned image searchable by magic. If pages[0].page_content is empty or garbled, add an OCR step and review the result before embedding it.

from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter

pages = PyPDFLoader("document.pdf").load()

splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=150,
    separators=["\n\n", "\n", ". ", " ", ""],
)
chunks = splitter.split_documents(pages)

for index, chunk in enumerate(chunks):
    chunk.metadata["chunk_id"] = f"chunk-{index:05d}"

The loader’s metadata travels with each chunk, including its source and page index where available. That link is what allows the final application to show the actual supporting passage.

Chunk sizes are tuning parameters, not universal truths. Tables, code listings, footnotes, and multi-column PDFs often need document-specific handling. Create a small set of questions with known page-level answers and use it to test changes.

Create a Persistent Chroma Collection

Use the maintained Chroma and Ollama integrations:

from langchain_chroma import Chroma
from langchain_ollama import OllamaEmbeddings

embeddings = OllamaEmbeddings(model="nomic-embed-text")

vector_store = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    collection_name="local_pdf_qa",
    persist_directory="./chroma_db",
)

Chroma’s default HNSW distance is squared L2, not cosine. Chroma also supports cosine and inner-product spaces, but changing the distance function is not automatically an improvement. Use a space supported by the embedding function and evaluate it on your own questions.

Do not interpret an arbitrary distance as a universal confidence percentage. Score direction and scale depend on the selected distance function and embedding model. Calibrate any refusal threshold against labeled retrieval examples.

Retrieve and Label the Exact Context

Retrieve a bounded set of chunks and give each one a temporary citation ID:

def retrieve_context(question: str, k: int = 5):
    docs = vector_store.similarity_search(question, k=k)
    sources = {}
    blocks = []

    for number, doc in enumerate(docs, start=1):
        citation = f"S{number}"
        page_index = doc.metadata.get("page")
        page = page_index + 1 if isinstance(page_index, int) else "unknown"
        source = doc.metadata.get("source", "document.pdf")
        chunk_id = doc.metadata.get("chunk_id", "unknown")

        sources[citation] = {
            "source": source,
            "page": page,
            "chunk_id": chunk_id,
            "text": doc.page_content,
        }
        blocks.append(
            f"[{citation}] source={source} page={page} chunk={chunk_id}\n"
            f"{doc.page_content}"
        )

    return "\n\n".join(blocks), sources

The model sees the same IDs that the application maps back to retrieved text. This is stronger than asking it to invent a filename or page number from memory.

Ask for a Grounded Answer

Use ChatOllama from langchain-ollama and make refusal explicit:

import re
from langchain_ollama import ChatOllama

llm = ChatOllama(model="llama3.1:8b", temperature=0)

SYSTEM_PROMPT = """Answer only from the numbered source blocks.
Every factual statement must cite at least one source ID such as [S1].
If the sources do not contain the answer, reply exactly:
Not found in the retrieved sources.
Do not use outside knowledge and do not invent source IDs."""

def answer_question(question: str):
    context, sources = retrieve_context(question)
    response = llm.invoke([
        ("system", SYSTEM_PROMPT),
        ("human", f"Question: {question}\n\nSources:\n{context}"),
    ])
    answer = str(response.content).strip()

    cited = set(re.findall(r"\[(S\d+)\]", answer))
    invalid = cited.difference(sources)

    if invalid:
        return {
            "answer": "Answer rejected: it cited a source that was not retrieved.",
            "sources": sources,
        }

    if answer != "Not found in the retrieved sources." and not cited:
        return {
            "answer": "Answer rejected: no retrieved source was cited.",
            "sources": sources,
        }

    return {"answer": answer, "sources": sources}

This validation checks that citations refer to real retrieved chunks. It does not prove that every sentence is entailed by those chunks. Display the cited text beside the answer and require human review for decisions where an error matters.

A validation function that searches sources for hard-coded phrases does not verify the generated answer. Verification must be tied to the current question, answer, and returned source chunks.

Test Retrieval Separately from Generation

When an answer is wrong, first ask whether the relevant source was retrieved.

Build a small evaluation file containing:

question
expected document
expected page or section
answerable: yes/no
required facts

For each question, record whether a top-k result includes the expected passage. If it does not, changing the prompt will not solve the retrieval problem. Try better extraction, different chunk boundaries, a different embedding model, metadata filters, or a carefully evaluated distance function.

Then score answer behavior:

  • Did every factual sentence cite a returned source?
  • Does each citation actually support the sentence?
  • Did the model refuse when the answer was absent?
  • Were page numbers displayed from metadata rather than generated by the model?
  • Did the answer distinguish direct text from inference?

Keep an “unanswerable” set. A system tested only on questions with obvious answers never demonstrates that it can refuse.

PDF Failure Modes to Expect

PDF is a presentation format, and extracted reading order can be wrong. Common problems include:

  • scanned pages with no text layer;
  • two-column pages interleaved in the wrong order;
  • repeated headers and footers polluting every chunk;
  • tables flattened into ambiguous prose;
  • diagrams whose meaning is not present in text; and
  • page labels that differ from zero-based PDF page indexes.

Inspect representative chunks before trusting retrieval. For high-stakes documents, preserve a link to a rendered page and let reviewers compare the extraction against the original visual.

Know What the Citations Prove

A cited answer shows which retrieved passage supported the response; it does not prove that retrieval found the best passage or that the model interpreted it correctly. This pipeline returns the exact context, rejects invented citation IDs, and allows “not found” when the context is insufficient. Read the cited page before relying on a consequential answer.

Sources