RAG guardrails in Python: follow a request through the code

Published 20 September 2026

ai-engineeringragpythonllm-evaluation

When my course assistant returns an answer, I want to trace the evidence and policy decisions behind it. Its Python response object keeps both alongside the generated text, so I can follow the same record through the UI and evaluation. Whether the configured verifier ran is part of that record too.

The interactive project article explains the layers and their purpose. Here, I follow the Python implementation through the parts that make those decisions. You can read the relevant code in the embedded gists, then open the surrounding implementation at the exact lines.

The excerpts come from repository commit 66fde26, and each gist is pinned to its own revision. They preserve the source apart from common leading indentation. Imports, enclosing classes and helper functions remain in the repository, so treat these as reading excerpts. To run the application, use the repository checkout below.

Return a decision that the UI and evaluator can inspect

LearningAssistant.answer() returns an AssistantResponse. Alongside the answer, it keeps the disposition, guard triggers, retrieved evidence and grounding result in one object. That gives the UI and evaluation code the same record to work with.

The response keeps the answer and its decision evidence together Python

View 01-response-contract.py on GitHub Gist

Source: AssistantResponse, pipeline.py lines 22-39.

The disposition enum has four values: answer, block, redirect and abstain. I keep them separate because a policy rejection needs a different investigation from a question the documents couldn’t answer. The distinction also lets the evaluation count unnecessary abstentions.

There is a second distinction inside this response: retrieved chunks and supporting chunks. Retrieval finds candidates. When the configured verifier runs, it identifies the subset it considers support for the answer. grounding_supported=None records that a result hasn’t been supplied. The local extractive mode can return citations without running a model verifier, so a citation alone doesn’t imply verification.

Decide whether the request needs a model classifier

The input guard starts with the configured rules. A blocking trigger returns before retrieval. When a model classifier is configured, the first pass leaves similarity out of the blocking decision. The following code uses similarity candidates to decide whether to call that classifier.

Similarity candidates feed the decision to run the classifier Python

View 02-classifier-routing.py on GitHub Gist

Source: classifier routing, pipeline.py lines 117-137. The preceding input check and early return show where an explicit policy match stops the request.

The not triggers condition matters when reading classifier_strategy="always": even that strategy calls the classifier only if the earlier guard pass left no triggers. With the ambiguous strategy, a similarity candidate or the ambiguity helper can request classification. This is conditional control flow, so a request doesn’t necessarily visit every box in the funnel.

The classifier result handling checks its label and configured confidence threshold. An unsupported label takes the abstention branch. A label in the policy’s blocking set takes the blocking branch. That reported confidence is an input to a decision rule, not a measured probability that the decision is correct.

Policy configuration lives in GuardrailPolicy. It holds the rule collections, blocking labels, allowed visibility and response messages. Keeping those choices together lets me adjust a policy and compare the resulting dispositions on the same cases.

Retrieve evidence within the configured scope

After the input check, the assistant passes its course and permitted visibility to the retriever. In the vector backend, the main query can be followed by a separate query for policy material. That second query matters for a course assistant: a passage explaining a concept and a passage defining permitted tutoring can both be useful.

The vector retriever can add policy context to the main evidence Python

View 03-scoped-retrieval.py on GitHub Gist

Source: VectorRetriever.search, vector.py lines 185-211. The caller passes the scope, and _chroma_where constructs the metadata filter.

The policy query has its own retrieval budget. It keeps the same course and visibility arguments, adds a source-type filter, then appends chunks that the main query hasn’t already returned. The final context can therefore contain more than the main query’s top_k chunks. That is worth accounting for when comparing context size and model cost.

This prototype takes scope from its configuration. In a service with several users, the retrieval layer would need to obtain permitted scope from authenticated identities and document permissions. The configured course filter shown here isn’t that identity system.

Back in the pipeline, retrieved text passes through the configured context checks. The evidence selector can then apply separate score thresholds to ordinary evidence and policy sources. These scores belong to the chosen retrieval setup. They don’t establish that a passage entails the answer.

An academic_integrity trigger takes a separate tutoring branch. It retrieves policy context and returns the configured tutoring message. The general answer-generation and entailment path below belongs to the other branch. Both paths subsequently reach the output guard.

Keep the model adapter explicit about its API contract

The generator receives a question and the selected chunks. Model access uses the Python openai package, with configurable model names for the different roles. A custom base URL selects the Chat Completions path. Without one, this adapter uses the Responses path.

Request structured output, validate it and retry invalid payloads Python

View 04-compatible-answer-adapter.py on GitHub Gist

Source: OpenAIAnswerGenerator._answer_payload, openai_models.py lines 91-131. The client configuration supplies the base URL, credentials, timeout and transport-retry limit.

For another OpenAI-compatible backend, I would check the specific features used here: Chat Completions, response_format={"type": "json_object"} and the requested parameters. Remote embeddings need a compatible embeddings endpoint as well. Using the same client package makes the integration familiar, but the backend still has to implement these calls.

On the chat path, the application parses the JSON and validates it locally. On the Responses path, it also supplies a strict JSON schema. The answer validator requires the expected fields, a Boolean answerability value and non-empty text. A structurally valid answer can still be unsupported by the documents, which is why the pipeline has a separate grounding decision.

The loop allows up to three attempts for an invalid answer payload and includes validation feedback in the next attempt. Transport retries are configured separately. One logical generation step can therefore involve more than one API request, something I’d include when measuring cost and latency.

Check support before selecting the final citations

When the entailment verifier is configured and there are retrieved chunks, the pipeline sends it the question, candidate answer and those chunks. The returned EntailmentResult includes support identifiers, unsupported claims and a confidence value. The application then applies its own acceptance conditions.

Validate the verifier result and retain only its supporting citations Python

View 05-grounding-decision.py on GitHub Gist

Source: grounding decision and citation selection, pipeline.py lines 274-324.

grounding_ok requires a successful verifier result, support marked true, sufficient reported confidence, no unsupported claims and at least one supporting chunk. It also rejects support IDs outside the retrieved set. If those conditions fail, the response abstains and preserves the diagnostic fields.

On this verified path, the citation list comes from the intersection of retrieved chunks and accepted support IDs. This prevents the application from presenting every retrieved passage as supporting evidence. It still depends on the verifier making a sound judgment. A confidence threshold and valid chunk IDs don’t prove semantic support on their own.

The output guard call checks the resulting text and citations. The response then records the disposition together with the evidence used to reach it. Its latency field measures elapsed time for the assistant call. It isn’t a breakdown of time spent in each individual layer.

Test the branch without asking a live model to cooperate

The citation-selection test supplies two chunks and a verifier double that approves only one. That makes the expected behavior precise: the response must cite the approved chunk and retain the verifier’s decision.

A controlled verifier result makes citation selection testable Python

View 06-citation-selection-test.py on GitHub Gist

Source: citation-selection test, test_pipeline.py lines 510-534.

The retriever, generator and verifier here are test doubles. The test checks how the application uses a supplied verification result. Evaluating whether a real verifier makes the right judgment requires reviewed examples and a separate model evaluation. I find it useful to keep those two questions separate when deciding why a check failed.

To run this test from the pinned repository version, start in a directory where you keep projects:

Check out the article's source revision and run its citation testBash
git clone https://github.com/zhittsova/llm-rag-guardrails-ai-course-assistant.git
cd llm-rag-guardrails-ai-course-assistant
git switch --detach 66fde26b258c0b5ab40fadecb2b1a658fa57a542
uv --directory guardrails-llm-deployment sync --frozen --dev
uv --directory guardrails-llm-deployment run pytest \
  tests/test_pipeline.py::test_entailment_keeps_only_citations_for_supporting_chunks

This test doesn’t call a model API. For the local application and review interfaces, follow the visual walkthrough.

Build the confusion matrix from expected and actual dispositions

The evaluator compares the expected disposition with the recorded one. It creates rows for expected labels and columns for actual labels, then calculates precision, recall and F1 for each of the four outcomes.

Disposition metrics come from an explicit four-class confusion matrix Python

View 07-disposition-metrics.py on GitHub Gist

Source: _behavior_summary, evaluation.py lines 401-450.

The denominators are visible here. predicted counts a column, support counts a row, and true_positives is their diagonal cell. Undefined precision or recall becomes zero. Macro-F1 averages the F1 values across all four enum labels, so an absent class contributes zero under this implementation. With the default rounding enabled, that average uses the already rounded per-class F1 values. The function also accepts round_metrics=False when unrounded values are needed.

These are disposition metrics. They don’t measure claim accuracy or the quality of an explanation. The offline LLM judge evaluates saved outputs separately. Its score needs validation against reviewed examples, as described in the evaluation guide and the project’s judge study discussion.

For a concrete way to read the implementation, start with the citation test, then follow supporting_chunks through the grounding branch into AssistantResponse. You can see exactly which behavior the test establishes and where a model judgment still needs its own evidence.