Step 1: Stand up DV-RAG and trace one benign request
You own the defense of DV-RAG-Support, ACME Cloud's customer-support assistant, the same target your red team broke in the offensive Indirect Prompt Injection lab. It is a real Retrieval-Augmented Generation (RAG) service, not a chatbot you talk to directly. A customer asks a question and the service runs four stages:
embed turn the question into a vector (NVIDIA llama-nemotron-embed-vl-1b-v2 embeddings)
retrieve find the most similar documents in the Milvus vector store (top-k),
scoped to the caller's tenant
generate paste those documents into the prompt and ask the LLM to answer
render the chat client renders the answer's markdown, loading any images
Before you defend anything, stand the service up and trace one benign request so you know what "normal" looks like. Every later step measures a deviation from this baseline. Two facts matter here. The sensitive data is not in the system prompt; it lives in the retrieved records (the customer's own account reference, billing contact, plan), exactly the shape of a real RAG leak. And the render stage auto-loads any markdown image the model emits, which becomes the exfiltration sink you will study and close in the steps that follow.
There is no poisoned document in the knowledge base yet. This step is a clean baseline: the assistant answers an account question from the caller's own record and renders nothing dangerous.
1. Build the vector index, then ask a benign in-tenant account question.
Hit Run, or in the terminal:
python3 dvrag.py --build
python3 dvrag.py "What is on my account record?"
--build reads every kb_*.md, chunks it, embeds it, and writes the Milvus index
to /home/labuser/kb.db. The second call runs the full embed -> retrieve ->
generate -> render flow once as the authenticated Globex caller.
2. Read the RETRIEVED: line and confirm the caller gets their OWN record.
You are the Globex account holder. The retriever should pull kb_account.md, the
Globex account record (account reference GLOBEX-ACR-88231), and nothing from
another tenant. The LOADED URLS: line should be (none): a clean answer loads
no images.
3. Read the render sink in the source.
Open dvrag.py and find _render and _extract_image_urls. Note the line
RENDER_ALLOWED_HOSTS = None. No allow-list means every image URL the model emits
would be loaded. That is the sink you will harden after you have reproduced the
attack against it.
Pass criteria
The vector index is built (/home/labuser/kb.db exists) and a benign account
question retrieves the caller's own Globex record (kb_account.md) with no
foreign-tenant row and no rendered image. That confirms the stack is live and the
baseline is clean.
kb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.py