How RAG turns documents into answers, and what to evaluate

Published 20 September 2026

ai-engineeringraginformation-retrievalllm-evaluation

Suppose someone asks a course assistant, “Can I submit the assignment on Monday?” The document collection contains a deadline, a late-submission policy and last year’s course guide. Finding text about assignments is fairly easy. Working out which passage answers this particular question takes more care.

That’s an illustrative example of the problem I wanted to explore with my AI course assistant. A useful response needs relevant evidence, the right policy and a decision about whether the assistant has enough information to answer. If the course or assignment is unclear, it may need clarification before retrieval can settle anything.

Retrieval-augmented generation, or RAG, gives a language model external material to use when producing a response. It can make the source collection available at request time, but it also gives us several decisions to inspect. What did we retrieve? What did the model actually receive? Which claims does that evidence support?

What retrieval changes

Lewis and colleagues’ 2020 RAG paper combined a language model with an external retrieval index. In a document assistant, the practical idea is to fetch relevant material and condition generation on it. Retrieving text for a request doesn’t itself update the generator’s weights. RAG systems can still include trained or fine-tuned components.

That separation is useful when documents change. You can update the corpus and index, then evaluate answers against the new snapshot. You still need to verify that ingestion completed, obsolete versions no longer appear where they shouldn’t, and caches respect the change. The presence of a new file on disk doesn’t prove that a live answer used it.

Fine-tuning can help adapt model behavior or a task-specific capability. A longer context window lets you supply more material. Those choices can coexist with retrieval. I would compare them using the questions the application needs to answer, the available evidence, and the cost of maintaining each approach.

Follow the evidence through the system

The two paths below describe a general document RAG system. They are a design guide, not a claim that my prototype implements every listed control.

Path Step What I would inspect
Preparing documents Parse and identify the source Text, tables, headings, document version and provenance
Preparing documents Split and index the material Whether each passage retains enough context to make sense
Answering a request Resolve scope and access Which sources this authenticated user may receive
Answering a request Retrieve candidate evidence Which passages match the question and why
Answering a request Select context Which candidates actually reach the generator
Answering a request Generate and check the response Whether its claims follow from the supplied evidence

Chunk boundaries matter because a condition and its exception may sit in adjacent paragraphs. A small chunk can be easy to retrieve but incomplete. A large one can preserve the explanation while carrying unrelated material. I would try representative questions against both and inspect the selected passages before choosing a chunk size.

Keep source identifiers and versions alongside the text. They make it possible to show a useful citation, remove an obsolete source and reproduce a result later. For tables, check that extraction preserves the relationship between headers and values. A perfectly retrieved row is of little use if parsing detached it from the column that defines it.

In a service with several users, access control belongs in trusted application and retrieval code. Restrict candidates using the authenticated identity before unauthorized text reaches the model. Apply the same boundary to caches and other derived stores. A model instruction asking it to keep secrets can’t enforce document permissions.

Search by wording, meaning, or both

A request may use an exact course code or a phrase copied from a policy. Lexical retrieval, such as BM25, is a useful baseline for that kind of matching. Dense retrieval compares learned representations and can help when the question and source use different wording. A hybrid approach combines candidates or rankings from both.

Test these choices on the same queries to see whether the more elaborate configuration earns its cost. The BEIR study compared retrieval approaches across varied tasks and found BM25 a robust baseline. Its results are evidence for testing across domains, not a ranking to apply unchanged to every new corpus.

For nonzero embeddings q and d, cosine similarity is:

cos(q, d) = (q · d) / (‖q‖ ‖d‖)

It compares vector direction. With unit-normalized embeddings, it equals their dot product. A score of 0.8 isn’t an 80% probability that the passage answers the question. Score distributions depend on the embedding model and data, so thresholds need calibration with the actual configuration. The cosine definition is precise about the normalization.

A reranker can examine the query and candidate passages more closely before context selection. It can reorder retrieved evidence, but it can’t recover a missing passage unless the workflow performs another retrieval. Query rewriting or an additional search may help a difficult question, at the cost of more work and another opportunity to change its meaning. Keep the original question in the trace so that drift is visible.

The context has to answer the question

Return to the illustrative deadline question. Suppose the current policy says:

Submissions close on Friday at 17:00. Students with an approved extension should follow the date in their extension notice.

The passage supports the ordinary deadline and the existence of an exception. It doesn’t establish whether this student has an extension, or what date an extension notice gives. “Monday is fine” would add information the passage doesn’t contain. A useful answer could state the ordinary rule and explain what additional information is needed.

Now suppose retrieval also returns last year’s guide with a different deadline. The system needs a way to distinguish versions and establish which one applies. Counting both passages as relevant won’t resolve the conflict. This is why I would keep evidence sufficiency separate from topical similarity.

More context also needs evaluation. The Lost in the Middle study showed that the location of relevant information affected performance in its tested models and tasks. That motivates a test on the model you actually use: does the answer change when useful evidence moves or distractors appear? It doesn’t establish that every current model has the same failure pattern.

What a citation should let you check

A citation should lead to the source and, ideally, the passage supporting the associated claim. A valid document ID establishes that the document exists. Assessing support requires reading what it says in context.

For the deadline example, citing the policy beside “Monday is fine” would make the unsupported claim look easier to trust. I would check the claim-to-passage relationship, whether the answer includes the exception accurately, and whether it leaves any substantive claims unsupported.

Source truth is a further question. A response can faithfully repeat an outdated or incorrect document. A grounding check against that document won’t discover the problem by itself. Provenance, ownership and document updates are part of the application design, with responsibilities outside the generator.

Retrieved text is also external input. It may contain instructions that attempt to change the model’s behavior. OWASP’s prompt-injection guidance describes this indirect route. Keep document content separate from application authority, restrict tool capabilities, and test attacks against the deployed boundaries. Retrieval doesn’t make an instruction trustworthy merely because it appears in a document.

Measure retrieval and answers separately

For a ranked list of distinct documents, let R be the labeled relevant documents and K the first k documents in the list:

Recall@k = |R ∩ K| / |R|

If three documents are labeled relevant and the top five contain two of them, document recall@5 is 2/3. Specify how you handle queries with no relevant documents, and report chunk-level and document-level metrics under different names. Several chunks from one document shouldn’t silently count as several relevant documents.

Rank-sensitive measures such as nDCG@k answer a different question: how good is the ordering, given the relevance labels? A retrieval evaluation needs both the labels and the unit being ranked. Incomplete labels can penalize a valid alternative source, so inspect disagreements before treating every mismatch as a retrieval failure.

For the generated response, these questions deserve separate checks:

Question A possible check Interpretation limit
Did the response address the request? Task-specific reviewed criteria A fluent response may leave part unanswered
Are its claims supported? Claim-to-evidence review or a validated verifier Support depends on the supplied source and the reviewer
Do citations support their associated claims? Citation support and coverage checks An existing document ID alone is insufficient
Did the system stop appropriately? Expected disposition versus observed disposition Abstaining on everything sacrifices useful answers

Frameworks such as RAGAs make separate retrieval and generation checks easier to run. A metric name alone doesn’t specify your evaluation: record its version, inputs, judge model and aggregation. For model-based scoring, review a sample of judgments against your own criteria before relying on the aggregate.

Compare the complete system with a simpler baseline on the same cases as well. For a diagnostic run, supplying known relevant evidence can help distinguish retrieval problems from problems using the evidence. That controlled experiment is different from measuring the deployed pipeline, so report it separately. Keep latency and cost beside quality, including additional searches and verification calls. The evaluation guide connects these choices to classification, generated answers and sequences of actions.

Where my guardrails project fits

My course assistant uses four dispositions: answer, block, redirect towards permitted tutoring help, and abstain when the available evidence doesn’t support an answer. Its Python implementation uses the openai package with OpenAI-compatible endpoints. The project article shows the configured pipeline and interactive examples, while the walkthrough covers the local UI.

In the recorded calibration run, the complete configuration made 392 correct disposition decisions on a balanced 400-case set. All eight errors were abstentions on answerable requests. That is a measured trade-off in the prototype, and the separate system holdout remains unopened. It isn’t an estimate of performance on another organisation’s documents or traffic.

For a different domain, I would start with questions that expose its actual document problems: versions that disagree, missing evidence, ambiguous scope and sources a particular user shouldn’t receive. Save the retrieved passages and the final answer together. When a response is wrong, that record lets you investigate whether the system found the wrong material, selected incomplete context or made a claim the evidence couldn’t support.