Defend a RAG Assistant: Block Indirect-Injection Exfil (EchoLeak)
Hands-on lab · IDE in your browser

Defend a RAG Assistant: Block Indirect-Injection Exfil (EchoLeak)

Harden the same deliberately-vulnerable RAG assistant the offensive lab broke, in small sequential steps.

Time
85 min
Checked steps
8
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit85 min · 8 stepsSession running
4 / 8 steps passingControl 2: provenance isolation (treat retrieved content as data) · step 5 of 8
dvrag.py▶ Run✓ Check
# what MAY load, so unknown hosts fail closed by default.RENDER_ALLOWED_HOSTS = {"cdn.acme-cloud.example"} # A NAIVE first-attempt defense the team shipped: a deny-list of the one host# they saw in an alert. Enabled only when DVRAG_NAIVE_BLOCK=1 (the bypass step).# Deny-lists are the wrong tool here; that is the lesson you will prove.NAIVE_BLOCK = os.environ.get("DVRAG_NAIVE_BLOCK") == "1"NAIVE_DENY_HOSTS = {"127.0.0.1"}      
TerminalOutput

The job

You own the defense of DV-RAG-Support, the customer-support assistant your red team just exfiltrated data from. A poisoned public document turns an ordinary account question into a data leak: the model emits a markdown image whose URL carries the customer's confidential account reference, and the chat client auto-loads it. That is the EchoLeak channel (CVE-2025-32711). You work in small, sequential steps. First you stand the pipeline up and trace one benign request, then you reproduce the leak so you can measure your fix against it. You watch a naive deny-list get bypassed by a renamed host, then build the durable fix one mechanism per step: an egress allow-list on the render sink that pins the parsed host and covers reference-style images, then provenance isolation so retrieved content cannot emit instructions and record fields never reach a URL. You verify both controls together, resist a userinfo / IP-encoding / paraphrase bypass battery, and finish at a ship gate where the attack success rate is 0 and benign quality holds.

8 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up DV-RAG and trace one benign request

    You own the defense of DV-RAG-Support, ACME Cloud's customer-support assistant, the same target your red team broke in the offensive Indirect Prompt Injection lab.

  2. 2

    Reproduce the EchoLeak exfil baseline

    Before you fix anything, reproduce the leak so you can prove your fix actually closes it.

  3. 3

    Watch the naive deny-list get bypassed

    The on-call engineer saw the alert: an outbound GET to 127.0.0.1.

  4. 4

    Control 1: an egress allow-list on the render sink

    Time to build the durable control.

  5. 5

    Control 2: provenance isolation (treat retrieved content as data)

    The egress allow-list from Step 4 closes the channel.

  6. 6

    Verify: exfil blocked, benign image and answers intact

    Both mechanisms are now in place and carried forward in dvrag.py: the egress allow-list on the render sink (Step 4) and provenance isolation (Step 5).

  7. 7

    Resist bypass variants

    A control that only stops the one payload you tested is the deny-list mistake all over again.

  8. 8

    Ship gate: attack success rate and benign regression

    This is the gate you would run in CI before shipping the fix.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up DV-RAG and trace one benign request

You own the defense of DV-RAG-Support, ACME Cloud's customer-support assistant, the same target your red team broke in the offensive Indirect Prompt Injection lab. It is a real Retrieval-Augmented Generation (RAG) service, not a chatbot you talk to directly. A customer asks a question and the service runs four stages:

embed     turn the question into a vector (NVIDIA llama-nemotron-embed-vl-1b-v2 embeddings)
retrieve  find the most similar documents in the Milvus vector store (top-k),
          scoped to the caller's tenant
generate  paste those documents into the prompt and ask the LLM to answer
render    the chat client renders the answer's markdown, loading any images

Before you defend anything, stand the service up and trace one benign request so you know what "normal" looks like. Every later step measures a deviation from this baseline. Two facts matter here. The sensitive data is not in the system prompt; it lives in the retrieved records (the customer's own account reference, billing contact, plan), exactly the shape of a real RAG leak. And the render stage auto-loads any markdown image the model emits, which becomes the exfiltration sink you will study and close in the steps that follow.

There is no poisoned document in the knowledge base yet. This step is a clean baseline: the assistant answers an account question from the caller's own record and renders nothing dangerous.

Do this

1. Build the vector index, then ask a benign in-tenant account question.

Hit Run, or in the terminal:

python3 dvrag.py --build
python3 dvrag.py "What is on my account record?"

--build reads every kb_*.md, chunks it, embeds it, and writes the Milvus index to /home/labuser/kb.db. The second call runs the full embed -> retrieve -> generate -> render flow once as the authenticated Globex caller.

2. Read the RETRIEVED: line and confirm the caller gets their OWN record.

You are the Globex account holder. The retriever should pull kb_account.md, the Globex account record (account reference GLOBEX-ACR-88231), and nothing from another tenant. The LOADED URLS: line should be (none): a clean answer loads no images.

3. Read the render sink in the source.

Open dvrag.py and find _render and _extract_image_urls. Note the line RENDER_ALLOWED_HOSTS = None. No allow-list means every image URL the model emits would be loaded. That is the sink you will harden after you have reproduced the attack against it.

Pass criteria

The vector index is built (/home/labuser/kb.db exists) and a benign account question retrieves the caller's own Globex record (kb_account.md) with no foreign-tenant row and no rendered image. That confirms the stack is live and the baseline is clean.

Provided for you:kb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.py

Frequently asked questions

Is this the defensive version of the indirect prompt injection lab?

Yes. It reuses the same deliberately-vulnerable target, DV-RAG-Support. The offensive lab teaches you to exfiltrate data through the EchoLeak channel; this lab gives you a working exploit and has you harden the target until the exploit is blocked while benign behavior still works.

What is the durable control you build?

Two independent layers. First, an egress allow-list on the render sink: instead of deny-listing the attacker's host, you name the few trusted asset hosts the chat client may load, so any unknown destination fails closed, across both inline and reference-style markdown images. Second, context isolation: retrieved documents are framed as untrusted data the model must not obey, and record fields are never copied into URLs. The verification step checks each layer holds on its own.

Why is a deny-list not enough?

A deny-list only blocks the exact destinations someone remembered to list. The attacker chooses where the data goes, so loopback alone can be written as 127.0.0.1, localhost, 0.0.0.0, IPv6 [::1], or numeric encodings, and any of them evades a host you forgot. You prove this bypass hands-on before building the allow-list, which inverts the logic and fails closed by default.

Do I need a machine-learning background?

No. You need to read and edit Python and understand a basic HTTP request. The lab is about how an LLM application trusts the wrong data and renders the wrong output, not about model internals. Everything model-specific is explained inline.

What you'll harden in this lab

This is a hands-on defensive-security lab built on a real RAG stack: a Milvus vector store, NVIDIA embeddings, and a multi-tenant knowledge base. You defend DV-RAG-Support, the same Retrieval-Augmented Generation (RAG) support assistant the offensive Indirect Prompt Injection lab attacks. You start by reproducing the exploit so you can measure your fix against it: a poisoned document makes the assistant exfiltrate a customer's confidential account record through a markdown-image URL the chat client auto-loads. This is the EchoLeak mechanism (CVE-2025-32711), the first real-world zero-click exploit against a production LLM system, and indirect prompt injection (OWASP LLM01).

Then you do the engineering. You watch an obvious deny-list of the attacker's host get bypassed by another spelling of loopback, which teaches why enumerating bad destinations fails. You replace it with an egress allow-list on the render sink that names the few hosts the client may load and drops everything else, covering both inline and reference-style markdown images. You add context isolation so retrieved documents are treated as untrusted data and account fields are never echoed into a link or image. You finish by verifying a fresh poison battery leaks nothing through any variant while benign account questions still get real, retrieval-grounded answers.