LLM security course: defend against the OWASP LLM Top 10
Eight hosted labs that harden a real RAG assistant and tool agent, then prove each control holds while benign traffic still passes.
Part of the AI Red Teaming Course.
What you'll build
- A four-control-point guard harness (input mediation, retrieval and context control, output mediation, action authorization) that stops a four-attack battery
- An egress allow-list on the render sink plus output redaction that close the EchoLeak exfil channel, wired to an attack-success-rate CI gate
- A RAG firewall that rejects directive-shaped ingestion and enforces a server-side tenant predicate the caller cannot widen
- A secret vault boundary, canary tripwire, and fail-closed decoding leak detector that keep a system-prompt secret unrecoverable
- A least-privilege tool agent with per-argument authorization and a human approval gate, and a signed, hash-pinned tool registry that refuses shadowing manifests
About this path
Defending an LLM application means putting controls between untrusted text and the actions the system can take, then proving those controls hold without breaking normal use. Because the model reads retrieved documents, tool descriptions, and prior outputs as instructions, a single guardrail is rarely enough; defense in depth wires an independent control at each point where untrusted data enters or leaves. The controls line up with the OWASP Top 10 for LLM Applications: prompt injection, insecure output handling, sensitive information disclosure, excessive agency and supply chain. These labs harden the same deliberately-vulnerable Retrieval-Augmented Generation (RAG) assistant and tool-using agent that the offensive labs attack, which is exactly the guardrail and human-oversight work weighted in NVIDIA NCP-AAI and in the Governance, Safety & Risk Management domain of Anthropic's CCAR-P. The controls are cataloged against the OWASP Top 10 for LLM Applications and MITRE ATLAS by name.
You begin with the shape of a complete defense: four independent control points around a RAG assistant (input mediation, retrieval and context control, output mediation, and action authorization), plus a guardrail layer wired to an attack-success-rate (ASR) CI gate that goes red the moment a change re-opens a hole. Then you lock down the RAG pipeline: an egress allow-list and provenance isolation that close the EchoLeak markdown-image exfil channel, a RAG firewall that rejects directive-shaped ingestion and enforces a server-side tenant predicate the caller cannot widen, and a secret vault boundary with a canary tripwire that keeps a system-prompt secret out of the model's context. The last stage hardens the agent: least-privilege tool scope with per-argument authorization and a human-in-the-loop approval gate, a signed and hash-pinned tool registry that refuses poisoned, rug-pulled, or shadowing manifests, and a contextual output mediator that encodes for each sink so cross-site scripting, server-side request forgery, SQL injection, and code execution all fail.
Every lab reproduces the exploit for you first, so you can start here without doing the offensive labs. All eight run in the browser against a sandboxed target with no GPU, take 80 to 90 minutes, and end on a ship gate where attack-success-rate is zero and benign traffic still passes clean through.
Who should join
- Python basics and comfort reading a short script
- Familiarity with RAG and tool-using agents; the offensive labs help but are not required, since each defense lab reproduces the exploit for you first
- The mindset that a fix must be verified: every lab re-runs the exploit to prove the control holds and benign traffic still passes
Every step is checked against the live environment. Progress saves between sessions.
Outline
8 labs · about 12 hours
1. Defense in depth and a CI gate
Wire four independent control points around a RAG assistant, then build a guardrail layer with an attack-success-rate gate that catches regressions.
2. Lock down the RAG pipeline
Close the EchoLeak exfil channel, build a RAG firewall for ingestion and tenant isolation, and isolate a system-prompt secret.
3. Harden the agent and its tools
Re-scope a tool agent to least privilege with a human approval gate, sign and pin the tool registry, and mediate every output sink.
- 6Defend Excessive Agency: Re-scope a Tool Agent to Least Privilege (AuthZ + Human Approval Gate)Hosted lab · 90 min · AdvancedPro
- 7Defend the Agent Supply Chain: Verify, Pin, and Capability-Gate Your Tool RegistryHosted lab · 90 min · AdvancedPro
- 8Defend: A Contextual Output Mediator for XSS, SSRF, SQLi, and RCEHosted lab · 90 min · AdvancedPro
Run all 8 labs with Preporato Pro, plus every other lab and practice test.
$29.99 per month or $290 per year. Cancel any time.
Frequently asked questions
No. Each defense lab stands the vulnerable pipeline up and reproduces the exploit for you before you harden it, so you see exactly what the control has to stop. Doing the matching offensive lab helps, but it is not a prerequisite.
Guardrails, human-in-the-loop validation and regulatory controls are the Governance, Safety & Risk Management domain of Anthropic's CCAR-P, and safety guardrails and responsible-AI practice are graded on NVIDIA NCP-AAI. These labs build those controls against a real target. They are cataloged against the OWASP Top 10 for LLM Applications and MITRE ATLAS by name.
Each lab re-runs the original exploit against your hardened build and checks that attack-success-rate falls to zero while benign requests still answer, query, and fetch correctly. Several labs finish with fresh, renamed, and paraphrased bypass variants to prove the control generalizes rather than pattern-matching one payload.
The labs are included in Preporato Pro ($29.99 per month or $290 per year), which also covers every practice test and every other lab on the platform. Individual lab pages show the full brief before you subscribe.