Build a RAG Firewall: Reject Poisoned Ingestion and Enforce Tenant Isolation
Hands-on lab · IDE in your browser

Build a RAG Firewall: Reject Poisoned Ingestion and Enforce Tenant Isolation

Defend the same multi-tenant RAG assistant the offensive labs attack, in small sequential steps.

Time
80 min
Checked steps
8
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit80 min · 8 stepsSession running
4 / 8 steps passingControl 1: a server-side tenant predicate the caller cannot widen · step 5 of 8
dvrag.py▶ Run✓ Check
def retrieve(question, k=4, tenant_scope=None):    """Embed the question and ANN top-k, scoped to the caller's tenant + public.         text/tenant/sensitivity/source.    """        
TerminalOutput

The job

You are the defender on DV-RAG-Support, ACME Cloud's multi-tenant customer-support assistant, the same target the offensive RAG labs attack. You work in small, sequential steps. First you stand the pipeline up and trace one benign in-tenant request so you know what normal looks like. Then you reproduce the two handed-to-you exploits one at a time: a poisoned document that wins semantic top-k for a broad class of account questions and steers the assistant's answer, and a caller-supplied tenant scope that reads Initech's confidential contract from the Milvus store. You watch a naive deny-list firewall get bypassed by a fresh payload, then build the durable control one mechanism per step: a server-side tenant predicate the caller cannot widen, and an ingestion screen that rejects directive-shaped documents before they are indexed. You verify both exploits are blocked with benign traffic intact, then prove fresh, paraphrased, and renamed variants are all blocked.

8 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up DV-RAG and trace one benign request

    You are the defender on DV-RAG-Support, ACME Cloud's multi-tenant customer-support assistant.

  2. 2

    Reproduce attack A: poisoned ingestion steers the answer

    You have two working exploits in hand.

  3. 3

    Reproduce attack B: a widened tenant scope reads another tenant

    Now reproduce the second exploit, an isolation failure that does not need any planted document.

  4. 4

    Watch a naive deny-list get bypassed by a fresh payload

    After the incident, the team shipped the obvious fix and called it done.

  5. 5

    Control 1: a server-side tenant predicate the caller cannot widen

    Time to build the durable control.

  6. 6

    Control 2: an ingestion screen that rejects poison before indexing

    The tenant predicate from Step 5 is carried forward in dvrag.py.

  7. 7

    Verify: both exploits blocked, benign traffic intact

    Both mechanisms are now in place: the ingestion screen in firewall.py and the server-side tenant predicate in dvrag.py, carried forward from Steps 5 and 6.

  8. 8

    Resist bypass: fresh, paraphrased, and renamed attacks all blocked

    A control that only stops the one payload you tested is the deny-list mistake all over again.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up DV-RAG and trace one benign request

You are the defender on DV-RAG-Support, ACME Cloud's multi-tenant customer-support assistant. It is a real Retrieval-Augmented Generation (RAG) service, not a chatbot you talk to directly. A customer asks a question and the service runs four stages:

embed     turn the question into a vector (NVIDIA llama-nemotron-embed-vl-1b-v2 embeddings)
retrieve  find the most similar documents in the Milvus vector store (top-k),
          filtered to the caller's tenant
generate  paste those documents into the prompt and ask the LLM to answer
render    the chat client renders the answer's markdown

Before you defend anything, stand the service up and trace one benign request so you know what "normal" looks like. Every later step measures a deviation from this baseline. Two facts matter here. The knowledge base is multi-tenant: each kb_*.md carries a tenant: in its frontmatter, and retrieve() is meant to scope results to the caller plus shared public docs. And there is no secret in the system prompt; the sensitive data lives in the retrieved records, exactly the shape of a real RAG leak.

Do this

1. Build the vector index, then ask a benign in-tenant account question.

Hit Run, or in the terminal:

python3 dvrag.py --build
python3 dvrag.py "What is on my account record?"

--build reads every kb_*.md, runs each through the ingestion firewall (firewall.scan_document, accept-all as shipped), chunks it, embeds it, and writes the Milvus index to /home/labuser/kb.db. The second call runs the full embed -> retrieve -> generate -> render flow once as the authenticated Globex caller.

2. Read the RETRIEVED: line and confirm the caller gets their OWN record.

You are the Globex account holder. The retriever should pull kb_account.md, the Globex account record (account reference GLOBEX-ACR-88231, billing contact, plan), and nothing from another tenant. That confirms in-tenant retrieval works cleanly before you start attacking it.

3. Read retrieve(), build_index(), and scan_document() in the source.

  • build_index() offers every document to firewall.scan_document() first; the shipped firewall (firewall.py) returns True for everything.
  • retrieve() builds its tenant filter from a tenant_scope argument. Note where that scope comes from; you will harden it in Step 5.

Pass criteria

The vector index is built (/home/labuser/kb.db exists) and a benign account question retrieves the caller's own Globex record (kb_account.md) and no other tenant's row. That confirms the stack is live and in-tenant retrieval is clean.

Provided for you:firewall.pykb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.py

Frequently asked questions

What will I actually build?

Two enforced controls on a real RAG service. An ingestion firewall (firewall.py) that rejects directive-shaped documents at index build time, so a poisoned document never enters the vector store, and a server-side tenant predicate in retrieve() that derives the scope from the authenticated session and validates tenant values against an allow-list, so a caller cannot widen access to another tenant.

Why is a deny-list not enough?

A deny-list of known-bad strings remembers one example payload while ignoring the injection pattern itself. The lab proves this: a fresh canary and a hostname sink sail past the literal deny-list, and a metadata-filter-injection value slips past a blocklisted tenant name. You replace both with controls that screen for the structure of the attack across payloads.

How does the lab confirm my control holds without breaking the assistant?

The grader plants a fresh poison every run, including a paraphrased variant with a different wording, and replays every cross-tenant scope payload, so you cannot pass by deleting an artifact or by tuning to one phrasing. It also checks that legitimate documents still index and that a benign same-tenant question still returns a real answer, so an over-blocking fix fails too.

What you'll do in this lab

This is a hands-on defensive-security lab on hardening a real Retrieval-Augmented Generation pipeline: a Milvus vector store, NVIDIA llama-nemotron-embed-vl-1b-v2 embeddings, chunking, and tenant-filtered top-k retrieval feeding a live LLM. You are handed two working exploits against DV-RAG-Support and your job is to build a RAG firewall that stops them. First you reproduce the baseline so you understand exactly what fires: a planted document that wins retrieval and steers the answer, and a widened tenant scope that reads another tenant's confidential contract.

Then you watch the obvious fix fail. A deny-list of the strings from the incident report is bypassed by a fresh canary and a hostname sink, the same way a blocklisted tenant value is bypassed by a metadata-filter-injection string. With that lesson in hand you build the durable control: an ingestion firewall that screens documents for the injection pattern (a self-declared answer policy, an imperative aimed at the model, a planted output sink) and rejects them before they reach the index, plus a retrieval predicate that derives the tenant scope server-side from the authenticated session and validates every tenant value against an allow-list. You finish by proving fresh and reworded exploits are both blocked while legitimate same-tenant answers still work.