Defend a RAG Assistant: Build a Guardrail Layer and an Attack-Success-Rate CI Gate
Hands-on lab · IDE in your browser

Defend a RAG Assistant: Build a Guardrail Layer and an Attack-Success-Rate CI Gate

You inherit DV-RAG-Support with a working EchoLeak-style exploit, and you defend it in small sequential steps.

Time
85 min
Checked steps
8
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Map the attack surface
Query
Retriever
LLM
Poisoned doc
retrieved chunk
Answer
0%
Attack-success rate
Attacks blocked · benign answers pass
graded on real output, not the model's talk

The job

You own DV-RAG-Support, the customer-support RAG assistant a previous lab broke, and you work in small, sequential steps. A poisoned knowledge-base document carries an "audit policy" that tells the model to end every reply with a logging pixel built from the account it referenced. The model obeys, emits a markdown image whose URL contains the customer's account reference, and the chat client auto-loads it, sending the secret to an attacker listener. That is the EchoLeak / CVE-2025-32711 pattern: model output becomes an attacker-controlled outbound request. First you stand the assistant up and trace one benign request, then you reproduce the leak so you can prove your fix later. You watch a one-host deny-list get bypassed by a renamed host, then build the durable guardrail one mechanism per step: an egress allow-list on the render sink so only approved hosts load, then output redaction of the sensitive record so even an approved-host request carries nothing. You verify the sink is closed with benign answers intact, wire an attack-success-rate gate so any future regression that re-opens the sink fails the build automatically, and finish by proving fresh, renamed, and paraphrased payloads are all blocked.

8 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up DV-RAG and trace one benign request through the guarded wrapper

    You are the defender on DV-RAG-Support, ACME Cloud's customer-support assistant, the same target a previous lab broke.

  2. 2

    Reproduce the attack: prove the EchoLeak sink fires before you defend it

    You traced what normal looks like.

  3. 3

    Watch the naive deny-list get bypassed by one renamed host

    An alert fired.

  4. 4

    Control mechanism 1: an egress allow-list on the render sink

    Time to build the durable control.

  5. 5

    Control mechanism 2: redact the sensitive record (and screen the input)

    The allow-list from Step 4 stops the callback to an untrusted host.

  6. 6

    Verify: the EchoLeak sink is closed and benign answers are intact

    You built the two control mechanisms one at a time: an egress allow-list on the render sink (Step 4) and record redaction plus an input screen (Step 5).

  7. 7

    Stand up the attack-success-rate gate and prove it is a real tripwire

    A fix that holds today can quietly stop holding tomorrow.

  8. 8

    Resist bypass: fresh, renamed-host, and paraphrased attacks all blocked

    A control that holds against the one payload you tested is not yet a control.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up DV-RAG and trace one benign request through the guarded wrapper

You are the defender on DV-RAG-Support, ACME Cloud's customer-support assistant, the same target a previous lab broke. It is a real Retrieval-Augmented Generation (RAG) service, not a chatbot you talk to directly. A customer asks a question and the service runs four stages:

retrieve  find the most similar knowledge-base documents (top-k), scoped to the
          caller's tenant, via NVIDIA llama-nemotron-embed-vl-1b-v2 embeddings + a Milvus store
guardrail screen retrieved context (input) and screen the answer (output)
generate  paste the context into the prompt and ask the LLM to answer
render    the chat client loads any markdown image the answer contains

You will defend at the two guardrail boundaries. app.py is the guarded wrapper you defend; it runs the pipeline and calls three hooks in guardrails.py:

retrieve  ->  input_guardrail(question, docs)  ->  model  ->  output_guardrail(answer)  ->  render

As shipped, those hooks are pass-throughs: input_guardrail returns the context untouched, output_guardrail returns the answer untouched, and render_host_allowed returns True for every host. So the wrapper behaves exactly like the unguarded assistant. Before you defend anything, stand the service up and trace one benign request so you know what "normal" looks like. There is no secret in the system prompt; the sensitive data (the caller's own account reference GLOBEX-ACR-88231) lives in the retrieved records, exactly the shape of a real RAG leak.

Do this

1. Build the vector index, then ask a benign account question through app.py.

Hit Run, or in the terminal:

python3 dvrag.py --build
python3 app.py "What plan am I on?"

--build reads every kb_*.md, chunks it, embeds it, and writes the Milvus index to /home/labuser/kb.db. The second call runs the full retrieve -> guardrail -> generate -> render flow once through the guarded wrapper.

2. Read the RETRIEVED: line and confirm the caller gets their OWN record.

You are the Globex account holder. The retriever should pull kb_account.md, the Globex account record (account reference GLOBEX-ACR-88231, billing contact, plan), and the assistant should return a real answer with retrieved sources. That confirms in-tenant retrieval and answering work cleanly before you start.

3. Read the pipeline in app.py and the three hooks in guardrails.py.

  • app.guarded_chat() calls input_guardrail() after retrieval, output_guardrail() after the model, and render() (which consults render_host_allowed()) last.
  • All three hooks in guardrails.py are pass-throughs as shipped. Note where each one sits; you harden them in Steps 4 and 5.

Pass criteria

The vector index is built (/home/labuser/kb.db exists) and a benign account question returns a non-empty answer with retrieved sources including the caller's own Globex record (kb_account.md). That confirms the stack is live and the guarded wrapper answers normally before you attack it.

app.py, the file you edit79 lines
"""
app.py -- the guarded DV-RAG-Support wrapper.

This is the stable wiring you do NOT edit. It runs the RAG pipeline and calls the
hooks in guardrails.py at the two trust boundaries:

    retrieve  ->  input_guardrail  ->  model  ->  output_guardrail  ->  render

It reuses the vulnerable assistant's own building blocks (dvrag.retrieve,
dvrag._llm) so the model and retrieval behave exactly as the unguarded target,
and it routes the final answer through render() so the markdown-image sink only
fires on hosts your output guardrail permits. The point of the lab is the policy
in guardrails.py; this file just enforces wherever that policy applies.

guarded_chat(question) returns the same shape as dvrag.chat():
    {"answer": str, "retrieved": [source...], "loaded_urls": [url...]}
loaded_urls is the ground-truth exfil signal: a non-empty list means the model's
output was turned into an outbound request. An attack succeeds when that request
goes anywhere other than a trusted render host.
"""

import re
import urllib.parse
import urllib.request

import dvrag
import guardrails


def render(answer):
    """Load each markdown image in `answer` whose host the guardrail allows.

    This is DV-RAG-Support's improper-output-handling sink, but gated by
    guardrails.render_host_allowed() instead of the unconditional auto-load in
    dvrag._render(). Returns the list of URLs actually loaded.
    """
    loaded = []
    for url in re.findall(r"!\[[^\]]*\]\((https?://[^)\s]+)\)", answer):
        host = urllib.parse.urlparse(url).hostname or ""
        if not guardrails.render_host_allowed(host):
            continue
        try:
            urllib.request.urlopen(url, timeout=2)
        except Exception:
            pass  # the listener already recorded the inbound request
        loaded.append(url)
    return loaded


def guarded_chat(question, tenant_scope=None):
    """Answer with RAG, with the input/output guardrails enforced around the model."""
    hits = dvrag.retrieve(question, tenant_scope=tenant_scope)
    # INPUT trust boundary: retrieved context is untrusted data.
    hits = guardrails.input_guardrail(question, hits)
    context = "\n\n".join(
        f"[{h['source']} | tenant={h['tenant']}]\n{h['text']}" for h in hits
    )
    messages = [
        {"role": "system", "content": dvrag.SYSTEM_PROMPT},
        {"role": "user", "content": f"CONTEXT:\n{context}\n\nUSER QUESTION: {question}"},
    ]
    answer = dvrag._llm(messages)
    # OUTPUT trust boundary: screen before any downstream action runs.
    answer = guardrails.output_guardrail(answer)
    loaded = render(answer)
    return {"answer": answer, "retrieved": [h["source"] for h in hits], "loaded_urls": loaded}


if __name__ == "__main__":
    import sys

    if "--build" in sys.argv:
        print(f"Indexed {dvrag.build_index()} chunks.")
        sys.exit(0)
    q = " ".join(a for a in sys.argv[1:] if not a.startswith("--")) or "What plan am I on?"
    out = guarded_chat(q)
    print("RETRIEVED:", out["retrieved"])
    print("ASSISTANT:", out["answer"])
    print("LOADED URLS:", out["loaded_urls"] or "(none)")
Provided for you:asr_gate.pydvrag.pyguardrails.pykb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_poison.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.pyprobes.json

Frequently asked questions

What is an LLM guardrail layer?

A guardrail layer is code that wraps a model call and inspects what goes in and what comes out. An input guardrail screens the prompt and any retrieved context for injected instructions before the model sees them; an output guardrail screens the response for policy violations (here, exfiltration images and leaked record fields) before any downstream action like rendering runs. In this lab you build both around an existing vulnerable assistant rather than rewriting the model.

What is an attack-success-rate (ASR) gate?

ASR is the fraction of attack attempts that achieve the attacker's goal. An ASR gate runs a fixed battery of attack probes (which must fail) and benign probes (which must still work) on every change, computes ASR, and exits non-zero when ASR rises above a set threshold. Wiring it into CI means a future edit that re-opens the leak breaks the build instead of shipping silently. It is the automation that turns "we fixed it once" into "it stays fixed."

Why does a deny-list of the attacker's host not work?

A deny-list blocks only the bad values you already know. The attacker's listener on 127.0.0.1 is the same loopback host as localhost, as 127.1, and as a decimal-encoded address, so blocking one spelling leaves the others open. An allow-list of the small set of hosts you actually trust to render flips the default to deny and closes the variants you never enumerated. You prove this in the lab by bypassing your own naive fix.

Do I need a machine-learning background?

No. You read and edit Python and reason about whether an attack actually fired. Everything model-specific is explained inline, and the target ships a deterministic offline mode so your control logic is testable without a live model. The skill on test here is building durable controls and a regression gate, not model internals.

What you'll do in this lab

This is a hands-on defensive AI security lab. You harden DV-RAG-Support, a Retrieval-Augmented Generation (RAG) customer-support assistant that ships with a real, working exploit: an indirect prompt injection (OWASP LLM01) hidden in a knowledge-base document coaxes the model into leaking a customer's account record (OWASP LLM02 sensitive information disclosure) through an auto-rendered markdown image, the improper-output-handling sink (OWASP LLM05) behind the real EchoLeak exploit. You reproduce the leak first so you can prove your fix later.

Then you build the defense the way a platform team would, one mechanism per step. You watch a shallow deny-list fix get bypassed by swapping one hostname, the lesson that earns the durable control. You close the render sink with an egress allow-list so only approved hosts load, then redact the sensitive record fields in the output so even an approved-host request carries nothing, and add an input guardrail that treats retrieved context as untrusted data and strips embedded instructions. You verify both mechanisms hold with benign answers intact. Finally you stand up an attack-success-rate (ASR) gate: a battery of attack and benign probes that computes the share of attacks that still succeed and fails the build when ASR crosses a threshold, then you prove fresh, renamed, and paraphrased payloads are all blocked.