Recon and Harness: Map a RAG Attack Surface and Measure Attack-Success-Rate
Hands-on lab · IDE in your browser

Recon and Harness: Map a RAG Attack Surface and Measure Attack-Success-Rate

Open your AI Red Team engagement against a real Retrieval-Augmented Generation assistant and build the methodology the whole path reuses.

Time
100 min
Checked steps
10
Level
Beginner
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit100 min · 10 stepsSession running
4 / 10 steps passingEncode one attack probe and fire it · step 5 of 10
probe.py▶ Run✓ Check
"""probe.py: encode ONE attack probe and run it against the target. Recon mapped the machine. Now you encode a single attack and see whether itfires. A probe in this lab has two halves, and naming them is the whole point:   DATA   the payload that rides into the model through untrusted retrieved         content. Here it is kb_poison.md, an "ACME audit and logging pixel"         policy that instructs the assistant to echo the retrieved account         reference into a markdown image URL aimed at the in-pod listener.   You do not type the payload at the model. You plant it in the knowledge base,   result["loaded_urls"]."""import os, sys, jsonos.chdir("/home/labuser"); sys.path.insert(0, "/home/labuser")import dvrag # The DATA half: the poison doc that must be present in the knowledge base.  POISON_DOC = "kb_poison.md" # The ACTION half: a benign account-record question a real customer would ask.   
TerminalOutput

The job

You are starting an engagement against DV-RAG-Support, a customer-support assistant backed by retrieval: a Milvus vector store, NVIDIA embeddings, and a multi-tenant knowledge base. You build the full methodology arc the rest of the path reuses, in small steps. You stand the service up and trace one request, then enumerate its endpoints, retrieval behavior, and output sink into a structured surface map. You encode a single probe and watch it fire, then build a deterministic side-channel oracle that counts real effects rather than the model's talk, and triage true positives from verbal-only false positives. You scale to an Attack-Success-Rate harness with a per-class breakdown, because a single successful trial against a stochastic target proves nothing, wire it into a CI gate that goes red on the vulnerable build, then apply the render-path allow-list and re-run the same gate to green while benign questions still answer.

10 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up DV-RAG and trace one request

    You are starting an engagement against DV-RAG-Support, ACME Cloud's customer-support assistant.

  2. 2

    Recon the retriever: probe top-k and write the first map block

    Recon is where every later finding comes from.

  3. 3

    Recon the render sink: confirm the EchoLeak auto-fetch channel

    The retriever decides what reaches the model.

  4. 4

    Finish the surface map: endpoints, tool hints, trust boundary

    Two blocks are mapped from real probes.

  5. 5

    Encode one attack probe and fire it

    Recon told you the surface.

  6. 6

    Build the deterministic success oracle

    You fired a probe.

  7. 7

    Triage true positives from false positives

    Your oracle is only useful if it counts the right things.

  8. 8

    Build the ASR harness: a battery with a per-class breakdown

    One probe, one run, proves nothing.

  9. 9

    Turn the harness into a CI gate (red on the vulnerable build)

    A number in a report is easy to ignore.

  10. 10

    Apply the fix and verify: gate green, benign still works

    The harness proved the vulnerability is real and reliable, and the gate is red.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up DV-RAG and trace one request

You are starting an engagement against DV-RAG-Support, ACME Cloud's customer-support assistant. It is a real Retrieval-Augmented Generation (RAG) service, not a chatbot you talk to directly. A customer asks a question and the service runs four stages:

embed     turn the question into a vector (NVIDIA llama-nemotron-embed embeddings)
retrieve  find the most similar documents in the Milvus vector store (top-k)
generate  paste those documents into the prompt and ask the LLM to answer
render    the chat client renders the answer's markdown, loading any images

Before you probe anything, stand the service up and trace a single benign request through those four stages so you know what "normal" looks like. Every later step measures a deviation from this baseline.

Do this

1. Build the vector index, then ask a normal account question.

Hit Run, or in the terminal:

python3 dvrag.py --build
python3 dvrag.py "What is on my account record?"

--build reads every kb_*.md, chunks it, embeds it, and writes the Milvus index to /home/labuser/kb.db. The second call runs the full embed -> retrieve -> generate -> render flow once.

2. Read the output and follow the flow in dvrag.py.

  • The RETRIEVED: line shows which documents the retriever pulled into the prompt. Confirm kb_account.md is among them: that is the customer's confidential account record (account reference GLOBEX-ACR-88231, billing contact, plan). That retrieved data is the exfil target later, exactly as the real EchoLeak exploit (CVE-2025-32711) exfiltrated a user's own data.
  • Read retrieve(), chat(), and _render() in dvrag.py. Note there is no secret in SYSTEM_PROMPT; the sensitive data lives in the retrieved record.

Pass criteria

The vector index is built (/home/labuser/kb.db exists) and an account-record question retrieves the customer's confidential account record (kb_account.md). That confirms the stack is live and that sensitive data flows into the prompt.

Provided for you:kb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.py

Frequently asked questions

Do I need to know machine learning to do this lab?

No. You need to read Python and understand a basic HTTP request. The lab is about how to map an LLM application's attack surface and how to measure an attack reliably, not about model internals. Everything model-specific is explained inline.

What is an Attack-Success-Rate (ASR) harness?

It is a test harness that runs the same attack many times against a target and reports how often it succeeds, as a fraction between 0 and 1. Because an LLM is non-deterministic, a single run tells you almost nothing. The harness fires a battery of probes across N trials, judges each with a deterministic oracle, logs every trial, and computes ASR = successful trials / total trials. You build a reusable one here and extend it through the rest of the path.

What is EchoLeak and why does the render sink matter?

EchoLeak (CVE-2025-32711) was a real zero-click exploit against Microsoft 365 Copilot. Hidden instructions made the assistant encode sensitive data into a markdown image URL, and the client auto-loaded that image, exfiltrating the data with no user interaction. The render sink in this lab is the same channel. Mapping it during recon and measuring how reliably it fires is exactly the skill this lab builds.

What you'll do in this lab

This is a hands-on offensive-security lab built on a real Retrieval-Augmented Generation (RAG) stack: a Milvus vector store, NVIDIA embeddings, and a multi-tenant knowledge base. You run the recon and measurement phases of an AI red-team engagement against a working support assistant called DV-RAG-Support, in small steps that build on each other. You stand the service up and trace one request, then enumerate its attack surface and emit a structured, machine-checkable surface map: the endpoints, the retrieval behavior, the markdown render sink, and the trust boundary the rest of the path attacks. Recon is where every later finding comes from, and you finish it with a deliverable a real engagement would hand off.

Then you build the instrument the whole path reuses. You encode a single probe and fire it, then build a deterministic side-channel oracle that counts a real effect, the confidential account reference leaving the pod through a URL the renderer actually loaded, rather than the model's talk. You triage that oracle against a benign case and a verbal-only false positive so you understand why a naive "did the model output bad text" detector over-counts. You scale to an Attack-Success-Rate (ASR) harness that fires a battery across N trials with a per-class breakdown and a structured JSONL audit log, because an LLM target is stochastic and one trial proves nothing. The behavior you measure is the EchoLeak markdown-image channel (CVE-2025-32711), the real-world zero-click pattern where a model encodes data into an image URL the client auto-loads. You wire the harness into a CI gate that goes red on the vulnerable build, then ship the render-path allow-list and watch the same gate go green while benign questions still answer, the methodology payoff that proves a fix actually works.