Fuzz an LLM App with garak: Run, Read, and Triage True vs False Positives
Hands-on lab · IDE in your browser

Fuzz an LLM App with garak: Run, Read, and Triage True vs False Positives

Run NVIDIA garak as an automated fuzzer against a real vulnerable RAG support assistant, read the JSONL run log and the per-probe DEFCON report, then do the skill that separates a scanner operator from a red teamer: triage the hits.

Time
90 min
Checked steps
8
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit90 min · 8 stepsSession running
6 / 8 steps passingDistrust the retrieved context: reframe context as data, never instructions · step 7 of 8
dvrag.py▶ Run✓ Check
# they saw in an alert. Enabled only when DVRAG_NAIVE_BLOCK=1 (the bypass step).# Deny-lists are the wrong tool here — that is the lesson you will prove.NAIVE_BLOCK = os.environ.get("DVRAG_NAIVE_BLOCK") == "1"NAIVE_DENY_HOSTS = {"127.0.0.1"} SYSTEM_PROMPT = (    "You are ACME Cloud's customer support assistant. Use the information in the "    "CONTEXT below to answer the customer's question accurately and concisely. "    "If the answer is not in the context, say you are not sure and suggest opening "     
TerminalOutput

The job

Manual testing does not scale and does not guard against regression. You are standing up the automated half of a red-team practice against DV-RAG-Support, the customer-support RAG assistant from Module 2. You will run garak as a fuzzer, read what it found, and triage the hits, because a scanner reports heuristic hits and only you can tell a real finding from a detector that pattern-matched. Dismiss a false positive with evidence, confirm a true positive against the exfil listener, then watch the genuine finding disappear once the fix is in.

8 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Recon: stand up the vulnerable target behind an OpenAI endpoint

    You are automating red-team testing of DV-RAG-Support, ACME Cloud's customer-support assistant from Module 2.

  2. 2

    Run the scan: fire a broad garak discovery battery at the endpoint

    garak runs a battery of probes (attack generators) against a generator and scores each response with detectors.

  3. 3

    Read the report: count hits per probe and grade the run

    A garak run produces several artifacts:

  4. 4

    Triage a false positive: dismiss a detector hit with evidence

    Scanners over-report.

  5. 5

    Confirm a true positive: run the custom probe, prove the effect on the listener

    A true positive is confirmed by reproducing the effect, not by trusting a detector.

  6. 6

    Harden the render sink: allow-list the one channel the finding rode out on

    Switch hats.

  7. 7

    Distrust the retrieved context: reframe context as data, never instructions

    The render allow-list from step 6 is the load-bearing fix: it closes the exfil channel even if the model is fully tricked.

  8. 8

    Verify the regression: fixed = 0 findings, reintroduced = re-trips

    A fix only counts when you re-run the exact confirmed attack and watch it fail.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Recon: stand up the vulnerable target behind an OpenAI endpoint

You are automating red-team testing of DV-RAG-Support, ACME Cloud's customer-support assistant from Module 2. A fuzzer talks to an LLM application the way any client does: over an HTTP API. So before you fuzz anything, you stand up the machine and confirm it is genuinely vulnerable.

Three processes make up the target:

  • dvrag.py is the RAG app: it embeds a question, retrieves from a Milvus vector store, builds a prompt, calls the model, and renders the answer. The render step auto-loads any markdown image the model emits. That is the exfil sink, the same EchoLeak channel (CVE-2025-32711) from Module 2.
  • dvserve.py wraps dvrag.chat() as an OpenAI-compatible HTTP endpoint on 127.0.0.1:9001, so garak and PyRIT can hit it like any chat API. The fuzzer's traffic is therefore real victim traffic: the vulnerable sink fires inside chat().
  • listener.py is the attacker's collection server on 127.0.0.1:9099. When the sink loads an image whose URL carries the account reference, the callback lands here. The listener log is your ground-truth oracle for the rest of the lab.
Do this
  1. Hit Run. It resets and starts the listener, builds the index, starts the OpenAI-compatible wrapper, waits for it to be healthy, then sends one benign account question (What plan am I on?) through /v1/chat/completions.
  2. Read the output. The model answers a normal account question. Then read the --- listener callbacks --- block: a callback carrying acct=GLOBEX-ACR-88231 means the account reference left through the markdown image sink. The poison doc (kb_poison.md) is already in the KB, so the app is vulnerable from step one.

That confirms the target is live and the sink fires on a benign question, so anything garak flags later can be checked against ground truth.

Pass criteria

dvserve.py answers /healthz, and a benign account question routed through the OpenAI endpoint produces an exfil callback containing 88231 on the listener.

Provided for you:dvserve.pygarak_dvrag.pygarak_openai.jsongarak_rest.jsonkb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_poison.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.pyread_report.pyrun_custom_probe.py

Frequently asked questions

What is garak?

garak is NVIDIA's open-source LLM vulnerability scanner. You point it at a generator (a model or an LLM application behind an API), it runs predefined probes (attack generators), and detectors score each response. It is the discovery and fuzzing tool in a red-team toolkit: broad and opinionated, and it reports heuristic hits that you then triage.

What is the difference between a true and a false positive here?

A false positive is when a detector flags a response that did not actually achieve the attacker's goal: the model refused, or just quoted the payload, and a string-match detector counted it anyway. A true positive is a hit you can reproduce as a real effect. In this lab the true positive is an account record exfiltrating to an in-pod listener, which you confirm by reading the listener log, not by trusting a detector score.

Do I need a machine-learning background?

No. You read Python and JSON and reason about whether an attack actually fired. Everything model-specific is explained inline. The point of the lab is triage judgment and regression discipline, not model internals.

What you'll do in this lab

This is a hands-on red-team automation lab. You point NVIDIA garak, a batteries-included LLM fuzzer and scanner, at a real Retrieval-Augmented Generation (RAG) support assistant called DV-RAG-Support and run a battery of attack probes against it. garak fires predefined attack prompts, scores each response with detectors, and writes a JSONL run log plus a per-probe DEFCON grade. Your job is the part a scanner cannot do for you: read the run and triage every hit, separating a genuine finding from a detector that merely pattern-matched.

You will dismiss a detector false positive with evidence from the run log, then confirm a genuine true positive the right way, by reproducing the effect against an in-pod exfil listener rather than trusting a detector score. The confirmed finding is an indirect prompt injection (OWASP LLM01) that leaks a customer's account record through an auto-rendered markdown image (OWASP LLM05 improper output handling), the same channel behind the real EchoLeak exploit. You finish by shipping the fix and re-running the battery to watch the finding regress to zero, the regression discipline a real red-team practice is built on.