AI red teaming course: break real systems, then ship the fix
Attack real LLM apps, RAG pipelines, and agents on live sandboxed targets, then ship and verify the fixes that stop them. For the people who build them and the people who break them.
What you'll break, and fix
Eight capabilities, previewed with each module's animation. Most labs end by writing and verifying the fix.
Map the AI attack surface. Where trusted instructions and untrusted data collide, mapped to OWASP and MITRE ATLAS.
Open this moduleAbout this course
Most AI security content is compliance PDFs and curse-word jailbreak demos. This is the opposite: a lab-first path where you attack deliberately-vulnerable LLM applications, RAG pipelines, and agents running in real sandboxes, and every exploit is confirmed by an automated check. You'll reproduce the techniques behind real-world incidents like EchoLeak and map each one to the OWASP LLM Top 10, the OWASP Agentic Top 10 (2026), and MITRE ATLAS. Then every module turns to defense, with a dedicated lecture, a hardening lab, and a project where you harden the same target and a grader re-runs your own exploit to prove the fix holds. It is also the closest thing to LLM penetration testing practice you can run in a browser: every module is a live engagement against a system you then harden, which is what AI security training has to look like to change how you build.
Skills you'll put on a resume
- Map the attack surface of an LLM app, RAG pipeline, and agent, and build a harness that measures attack-success-rate
- Exploit indirect prompt injection delivered through retrieved documents, tool output, and inter-agent messages
- Poison a RAG knowledge base and exploit retrieval, embedding, and cross-tenant isolation failures
- Exploit improper output handling into real sinks: XSS, SSRF, command and SQL injection, and unreviewed code execution
- Exploit excessive agency (confused deputy, tool-scope escalation, memory poisoning) and re-scope agents to least privilege
- Attack the agentic supply chain: MCP tool poisoning, rug pulls, tool shadowing, and inter-agent propagation
- Extract hidden system prompts and force sensitive-information disclosure
- Automate red teaming with garak, PyRIT, and promptfoo, and gate CI on attack-success-rate regression
- Run a full LLM/agent VAPT engagement and deliver a professional report with working exploits and verified remediations
For
For builders and breakers of LLM apps. If you ship RAG pipelines, assistants, or agents, you'll learn to attack your own system and harden it before someone else does. If you come from security, you'll learn how LLM apps are wired and where they break. We build both on-ramps
Prerequisites
- Comfortable with Python and calling an HTTP API (scripting a small client, reading a codebase)
- Curiosity about how LLM apps work or how they break. We build both the AI and the web-security mental models as we go
- No machine-learning background and no security background assumed
Every lab in this course, module by module
What you break, then what you fix
01AI Attack Surface & Pentest Methodology
An LLM application collapses the boundary that web security depends on: instructions and untrusted data travel in the same channel, so data can become commands. This module builds the mental model and the method. You'll learn the trust boundary that breaks, overlay the OWASP LLM Top 10, the OWASP Agentic Top 10, and MITRE ATLAS on a real agent architecture, scope an AI pentest (authorization, rules of engagement, testing a non-deterministic target), and build the repeatable harness and attack-success-rate metric you'll use for the rest of the path.
Read first: AI Penetration Testing: A Practical Guide to Testing AI Systems, LLM Pentesting Hands-On: Exploit a Live Model, Then Fix It
- Recon and Harness: Map a RAG Attack Surface and Measure Attack-Success-Rate100 min · beginnerOpen your AI Red Team engagement against a real Retrieval-Augmented Generation assistant and build the methodology the whole path reuses. Stand up the service and trace one request, enumerate its attack surface into a structured, machine-checkable map, encode a single probe, then build a deterministic side-channel oracle that counts real effects instead of the model's talk. Triage true positives from verbal-only false positives, scale to an Attack-Success-Rate harness with a per-class breakdown, wire it into a CI gate that goes red on the vulnerable build, then ship the render-path allow-list and watch the same gate go green while benign questions still answer.
- Defense in Depth: Wire Four Control Points Around a RAG Assistant90 min · advancedHarden DV-RAG-Support, a real Retrieval-Augmented Generation assistant, by building a guard harness with four independent control points one mechanism per step: input mediation, retrieval and context control, output mediation, and action authorization. You are handed a working four-attack battery (direct injection, cross-tenant retrieval, sensitive-field exfiltration, and an unauthorized image fetch). Stand the pipeline up, reproduce all four attacks one per stage, watch a single naive filter get bypassed, then build each control point in its own step so every attack class is stopped at its matching layer while a benign customer request still passes clean through all four. Verify the coverage matrix reads four-for-four, then prove reworded and renamed bypass variants are all resisted.
- Instrument the Four Control Points and Prove Coveragegraded project
02Indirect Prompt Injection
Direct injection is the curse-word demo. The real-world exploit is indirect: the payload arrives through data the model consumes, a retrieved document, a scraped page, a tool result, another agent's message. The victim never sees it, which is what made EchoLeak (CVE-2025-32711) a zero-click attack on Microsoft 365 Copilot. You'll plant a document in a support assistant's knowledge base, turn an innocent question into data exfiltration through an auto-loaded image, then close the channel.
Read first: What Is Prompt Injection? A Plain-English Guide with Real Examples, Prompt Injection Attacks Explained: How Attackers Hijack LLMs (OWASP LLM01), Indirect Prompt Injection: How AI Agents Get Hacked by the Content They Read
- Indirect Prompt Injection: Exfiltrate Data from a RAG Assistant75 min · advancedAttack a real Retrieval-Augmented Generation assistant end to end: a Milvus vector store, NVIDIA embeddings, and a multi-tenant knowledge base. Win semantic retrieval with a poisoned document, exfiltrate a customer's confidential account record through the EchoLeak markdown-image channel, measure attack-success-rate, bypass a naive defense, then ship the real fix.
- Defend a RAG Assistant: Block Indirect-Injection Exfil (EchoLeak)85 min · advancedHarden the same deliberately-vulnerable RAG assistant the offensive lab broke, in small sequential steps. Stand up the pipeline and trace one benign request, reproduce the EchoLeak markdown-image exfil, then watch a naive deny-list get bypassed by a renamed host. Build the durable fix one mechanism per step: an egress allow-list on the render sink that pins the parsed host (defeating userinfo, IP-encoded, and IPv6 spellings) and covers reference-style images, then provenance isolation so retrieved documents cannot emit instructions. Verify both controls together, resist a userinfo / IP-encoding / paraphrase bypass battery, and pass a final ship gate where attack success rate is 0 and benign quality holds.
- Write Up an Indirect Prompt-Injection Findinggraded project
- Build the Egress and Provenance Mediator that Closes the EchoLeak Channelgraded project
03RAG Pipeline Exploitation
RAG is the most common enterprise LLM pattern and its own attack surface: ingestion, embedding, retrieval, and tenant isolation each fail in distinct ways. You'll craft adversarial documents that win retrieval for a broad class of queries, exploit metadata-filter isolation bugs to read another tenant's data, and weaponize the retrieval pipeline to harvest secrets that were accidentally indexed (the ATLAS RAG-poisoning and credential-harvesting techniques).
- Retrieval Poisoning: Win Top-k Across a Whole Query Class and Steer the Answer80 min · advancedAttack a real Retrieval-Augmented Generation assistant where it is most exposed: retrieval. Plant one document in a Milvus + NVIDIA embeddings knowledge base, craft it to win cosine top-k for one account question, then for the whole account-query class, then steer the generated answer through a directive framed as routine policy. Measure broad-class coverage and steering attack-success-rate, then harden in two distinct moves: treat retrieved context as data behind a non-spoofable boundary, and cap how many top-k slots any single source may take. Re-run the same battery and watch attack-success-rate collapse.
- Cross-Tenant Leakage: Break RAG Metadata Isolation and Exfiltrate Another Tenant's Contract70 min · advancedAttack the multi-tenant isolation of a real Retrieval-Augmented Generation assistant. Two stacked bugs in one retriever, a caller-controlled tenant scope and a string-concatenated metadata filter, let a Globex-scoped caller read Initech's confidential contract from a Milvus + NVIDIA embeddings store. Chain the cross-tenant read into the EchoLeak markdown-image sink to exfiltrate the data to a listener, then harden the pipeline so isolation and the sink both hold.
- Build a RAG Firewall: Reject Poisoned Ingestion and Enforce Tenant Isolation80 min · advancedDefend the same multi-tenant RAG assistant the offensive labs attack, in small sequential steps. Stand up the pipeline and trace one benign request, then reproduce two handed-to-you exploits one at a time: a poisoned document that wins retrieval and steers the answer, and a caller-controlled tenant scope that reads another tenant's confidential contract. Watch a naive deny-list get bypassed by a fresh payload, then build the durable control one mechanism per step: a server-side tenant predicate the caller cannot widen, then an ingestion screen that rejects directive-shaped documents before indexing. Verify both exploits are blocked with benign traffic intact, then prove fresh, paraphrased, and renamed variants are all blocked on a real Milvus + NVIDIA embeddings stack.
- Prove a Knowledge-Base Exfiltration Chain End to Endgraded project
- Build a RAG Ingestion and Tenant-Isolation Firewallgraded project
04Improper Output Handling
The model's output flows into downstream sinks: a browser, a shell, a SQL query, another API. If the app trusts that output, the model becomes an injection vector into those sinks. You'll drive model output into stored XSS, pixel-exfiltration, SSRF, command execution, and unreviewed code execution (the OWASP Agentic ASI05 RCE class), then write the encode/validate/allow-list fix that blocks your own proof-of-concept.
Read first: EchoLeak Explained (CVE-2025-32711): The First Zero-Click AI Data Exfiltration Exploit
- Insecure Output Handling: Zero-Click Exfiltration Through Rendered Model Output (EchoLeak)85 min · advancedTreat the model as an untrusted source whose output flows into a sink: the chat client's markdown renderer. Prove the renderer auto-fetches, plant a document so a benign account question makes the assistant echo a customer's own account reference into a markdown image URL, and watch the renderer auto-fetch it (zero-click exfil, the EchoLeak channel). Measure its attack-success-rate, defeat a CSP-style allow-list through a first-party open proxy, measure that bypass too, then harden in two moves: close the render sink so untrusted output never fires an outbound request, add an audited host allow-list the open proxy cannot defeat, and re-run both attacks to watch ASR fall to zero.
- Insecure Output Handling: SSRF, SQLi, and Command Execution Through an Agent's Tools80 min · advancedRed-team OpsBot, a ReAct tool-using support agent, by shaping the arguments it passes to its own tools. File a poisoned support ticket and a benign on-call query turns into a server-side request forgery against a metadata endpoint, a SQL injection that drops a canary and reads another tenant's rows, and code execution through a transform helper. Measure attack-success-rate, then close every sink: an allow-listed fetch, parameterized queries, and a removed code tool.
- Defend: A Contextual Output Mediator for XSS, SSRF, SQLi, and RCE90 min · advancedDefend DV-ToolAgent, a tool-using support agent whose model output flows raw into four interpreters, in small sequential steps. Stand up the agent and trace one benign request, then reproduce each sink one at a time: a behavioral SSRF (the agent fetches an internal node-health endpoint), a behavioral cross-tenant SQL read (an entitlement comparison pulls another tenant's record), then OS command execution, a SQL write, and stored XSS demonstrated structurally at the sink. Watch a naive blocklist get bypassed by a variant, then build one contextual output mediator one sink-class at a time: an SSRF host allow-list with an IP-literal guard and parameterized tenant-scoped SQL, then refuse-arbitrary-code for the runtime and contextual HTML encoding for the browser. Verify every freshly planted payload and its bypass variant is blocked at the sink while benign output still renders, queries, fetches, and computes.
- Chain an Output-Handling Flaw to a Real Sink, Then Ship the Fixgraded project
- Build a Contextual Output-Encoding and Allow-List Mediatorgraded project
05Excessive Agency & Agentic Exploitation
Agents act. Over-broad tools, missing authorization, and trusted inter-agent messages turn a prompt-level foothold into real-world actions: the confused deputy. You'll escalate tool scope to perform actions beyond a user's authority, poison an agent's long-term memory so a payload re-triggers weeks later, then re-scope the agent to least privilege and prove the exploit is dead.
Read first: AI Agent Security: Excessive Agency, Confused Deputies and How to Contain an Agent
- Excessive Agency: Turn a Support Ticket into a Privileged Action (Confused Deputy)75 min · advancedAttack a real tool-using ReAct agent end to end. As the low-privilege ticket-ingest account, plant an authorized-looking record correction in a support ticket and make DV-ToolAgent run a privileged billing-payee redirect under its own shared credential (the confused deputy), then reach an internal-only endpoint and exfiltrate its value through the fetch tool (SSRF via tool args). Measure attack-success-rate, then harden the tool boundary, scope the DB tool read-only, carry per-user authorization, and allow-list the fetch tool, and prove your own exploit is dead while normal lookups still work.
- Memory Poisoning: Plant a Note That Re-Fires in a Fresh Session (Persistence)80 min · advancedAttack the long-term memory of a real tool-using ReAct agent. As the low-privilege ticket-ingest account, plant a single benign-looking routing note into DV-ToolAgent's shared, un-namespaced memory store through an ingested ticket. In a brand-new session for a different legitimate user, the agent recalls your note and redirects a GLOBEX invoice to your payee, persistence across the session boundary that single-turn filters never see. Use MINJA-style progressive shortening so the stored record reads as a mundane preference, measure attack-success-rate against benign controls, then harden memory with per-user namespacing and a data-only quarantine and prove the poison dead while legitimate recall still works.
- Defend Excessive Agency: Re-scope a Tool Agent to Least Privilege (AuthZ + Human Approval Gate)90 min · advancedHarden DV-ToolAgent, a real tool-using ReAct agent, against the confused-deputy and scope-escalation attacks from the offensive lab, in small sequential steps. Stand up the agent and trace one benign ticket, then reproduce the handed-to-you exploit one surface at a time: an ingested ticket that makes the agent redirect a billing payee under its own shared credential and read across tenants, then an SSRF reach into an internal-only endpoint and a poisoned-memory replant. Watch a naive SQL denylist get bypassed by a case-folded variant, then build the durable control one mechanism per step: least-privilege tool scope (the ingest role holds no write scope), per-argument authorization decided on the session identity (a caller-supplied identity claim is ignored), and a human-in-the-loop approval gate that holds high-impact writes pending an explicit token. Verify the exploit is dead with authorized in-scope work intact, then prove obfuscated, renamed, and spoofed variants are all blocked.
- Exploit an Over-Privileged Agent, Then Re-Scope It to Least Privilegegraded project
- Re-Scope a Tool Agent to Least Privilege with an Approval Gategraded project
06Agentic Supply Chain & Multi-Agent Exploitation
The hottest 2026 surface. An agent trusts a tool's description as much as its code, so a poisoned tool description hijacks it. You'll poison an MCP tool, pull a rug (mutate a tool after approval), shadow a trusted tool with a malicious twin, and inject a payload that propagates from one agent to another (the Morris II self-replicating pattern). Then you'll harden the tool manifest so the poison is rejected.
Read first: MCP Security Best Practices: Tool Poisoning, Token Passthrough and the Controls That Hold, Did Escaped AI Agents Leave Self-Replicating Code Across the Internet? Tracing the Claim
- MCP Tool Poisoning: Hijack an Agent Through a Tool Description (and a Rug Pull)95 min · advancedAttack a real MCP-style tool registry end to end. OrbitDesk's support agent loads its tools from a runtime registry and reads each tool description as trusted instruction text. Register a poisoned tool whose description hides a routine-looking audit directive, make the agent read an account record and forward its reference to your in-pod collector (data exfiltration through tool metadata), then pull a rug: register the tool benign, pass review, and silently mutate its description after approval. Measure attack-success-rate, then harden the registry, scan descriptions for hidden instructions and pin approved tool objects, and prove a fresh poison and a fresh rug pull are both dead while benign tickets still resolve.
- Tool Shadowing: Hijack an Agent's Tool Selection With a Name Collision75 min · advancedAttack a real MCP-style tool registry by shadowing a trusted tool. OrbitDesk's support agent loads its tools from a runtime registry with a flat namespace and last-write-wins resolution, and it picks which tool to call from attacker-controllable descriptions. Register a malicious twin with the same name as the trusted record reader and a more compelling, compliance-approved description, so the agent calls your twin instead. The twin silently reads a cross-tenant record and exfiltrates its reference, and the same shadow hands another tenant's data straight back to the caller. Measure how reliably the shadow wins selection, then harden the registry, namespace tools and reject collisions, and prove a fresh shadow is dead while benign lookups still work.
- Inter-Agent Injection: Propagate a Morris II Worm Across a Two-Agent Graph85 min · advancedAttack a real two-agent support graph where one agent's output is the next agent's input with no authentication and no validation. Plant a self-replicating payload in the only untrusted input, an inbound customer email the Intake agent ingests. Intake forwards it, and the Resolver agent, trusting the inter-agent notes as instructions, performs the attacker-directed action AND re-emits the payload verbatim: a second-hop cascade and the Morris II replication primitive, bounded to two hops. The harness attributes every side effect to the agent that caused it, so you prove the second agent executed. Measure propagation reliability, then harden the channel, a schema-constrained, validated, replication-aware handoff, and prove the cascade is contained while benign tickets still resolve.
- Defend the Agent Supply Chain: Verify, Pin, and Capability-Gate Your Tool Registry90 min · advancedHarden a real MCP-style tool registry until a poisoned, rug-pulled, or shadowing tool manifest cannot reach the agent, in small single-concept steps. OrbitDesk's support agent loads its tools from a runtime registry and reads each tool description as trusted instruction text. You stand the registry up and trace one benign ticket, then reproduce three techniques one at a time: a poisoned tool description that hijacks the agent into leaking an account reference, a post-approval rug-pull mutation the registry serves with no pin, and a shadowing twin the flat namespace selects under a trusted tool name. You apply the obvious fix, a description blocklist, and watch a clean-description variant defeat it through its delegate. Then you build the durable control one mechanism per step: manifest signature verification, then hash pinning (rug-pull / change detection), then a per-tool capability allow-list with namespacing. You finish by proving that freshly planted unsigned, forged, mutated, and shadowing manifests are all refused while a legitimate signed and pinned tool stays admitted and usable, and an inter-agent worm's second hop is contained.
- Audit an Agent's Tool Supply Chain and Multi-Agent Graphgraded project
- Build an MCP Manifest Verifier with Signing, Pinning, and an Allow-Listgraded project
07Model & System-Prompt Leakage
System prompts, embedded secrets, and training data leak through the model. The system prompt is not a secret store, and proving it is a core finding. You'll extract a hidden system prompt and its tool definitions, then force the assistant to disclose an embedded secret and another user's data reachable through over-broad context.
- System-Prompt Extraction: Recover a RAG Assistant's Hidden Instructions70 min · intermediateRed-team Aria, a real Retrieval-Augmented Generation support assistant: a Milvus vector store, NVIDIA embeddings, and an LLM that answers from one shared context window. Confirm a hidden system prompt with an embedded secret exists, recover it through ordinary chat by direct echo, climb the extraction ladder when a stronger refusal posture resists, defeat a naive output filter with encoding-egress, measure the extraction Attack-Success-Rate across the techniques, then ship the durable fix (minimize the secret out of the prompt) and verify extraction yields nothing.
- Sensitive Data Disclosure: Leak Confidential Records from a RAG Assistant75 min · intermediateAttack a real Retrieval-Augmented Generation assistant where the system prompt only asks for privacy: a Milvus vector store, NVIDIA embeddings, and a multi-tenant knowledge base. Force disclosure of your own gated billing secret, pull another customer's record across a disabled tenant filter, harvest an accidentally-indexed service key, then ship the real fix with pre-retrieval authorization, corpus hygiene, and output redaction.
- Defend: Secret Isolation for a RAG Assistant80 min · advancedHarden the same RAG support assistant that the extraction lab broke, in small sequential steps. A live signing key, an internal build id, and a canary token are baked into the system prompt, so the secret is exposed by construction: it shares one context window with the customer's message. Stand the service up, reproduce the exposure (the secret is present in the model's context), and watch a naive cleartext output filter fall to encoding-egress. Then build the durable control one mechanism per step: a vault boundary that holds the secret out of the model's context, a seeded canary tripwire, and a fail-closed decoding leak detector that matches the secret and its Base64/ROT13/hex forms. Verify the secret is unrecoverable and benign answers are intact, then resist a re-planted, re-encoded exfil battery.
- Isolate Secrets Behind a Tool Boundary and Add Canary Leak Detectiongraded project
08Automation, Fuzzing & Red-Team Tooling
Manual testing doesn't scale or guard against regression. A real AI red-team practice is automated, measured by attack-success-rate, and wired into CI. You'll point garak and PyRIT at a target, triage true from false positives, and build a harness that runs a battery of attacks (including the sophisticated multi-turn jailbreaks: Crescendo, Skeleton Key, Policy Puppetry) and fails CI when attack-success-rate crosses a threshold.
- Fuzz an LLM App with garak: Run, Read, and Triage True vs False Positives90 min · advancedRun NVIDIA garak as an automated fuzzer against a real vulnerable RAG support assistant, read the JSONL run log and the per-probe DEFCON report, then do the skill that separates a scanner operator from a red teamer: triage the hits. Dismiss a detector false positive with evidence, confirm a genuine indirect prompt injection against an in-pod exfil listener, and watch the finding regress to zero after the fix ships.
- Defend a RAG Assistant: Build a Guardrail Layer and an Attack-Success-Rate CI Gate85 min · advancedYou inherit DV-RAG-Support with a working EchoLeak-style exploit, and you defend it in small sequential steps. Stand the assistant up and trace one benign request, then reproduce the leak: an indirect prompt injection makes the model echo a customer's account record into a markdown image, and the renderer fires it as an outbound request that exfiltrates the data. Watch a naive host deny-list get bypassed by a renamed host, then build the durable guardrail one mechanism per step: an egress allow-list on the render sink so only approved hosts load, then output redaction of the sensitive record so even an approved-host request carries nothing. Verify the sink is closed with benign answers intact, then stand up an attack-success-rate (ASR) gate that runs a probe battery on every change: green while guardrails hold, non-zero the moment a regression re-opens the sink, green again when you back it out. Finish by proving fresh, renamed, and paraphrased payloads are all blocked.
- Build a CI Red-Team Harness That Gates on Attack-Success-Rategraded project
- Build a Guardrail Layer and an Attack-Success-Rate CI Gategraded project
Lab paths through this course
The same labs, grouped by topic
- Prompt injection examples you can runFourteen worked prompt injection examples against a sandboxed RAG assistant and tool agent, each measured by attack success rate.
- LLM security course: defend against the OWASP LLM Top 10Eight hosted labs that harden a real RAG assistant and tool agent, then prove each control holds while benign traffic still passes.
Guides & articles
Deep-dive reading that pairs with this course
What Is Prompt Injection? A Plain-English Guide with Real Examples
Prompt injection explained from zero: what it is, why LLMs fall for it, five real examples walked through step by step, and what actually stops it. The beginner's guide to OWASP's number one LLM risk.
ReadPrompt Injection Attacks Explained: How Attackers Hijack LLMs (OWASP LLM01)
Prompt injection is the number one LLM security risk (OWASP LLM01). Learn how direct and indirect attacks hijack AI assistants, the real CVEs (EchoLeak), what we found testing it against a live model, and how to actually defend.
ReadIndirect Prompt Injection: How AI Agents Get Hacked by the Content They Read
Indirect prompt injection hides attacker instructions in documents, emails, and web pages your AI reads. A deep dive into the kill chain, the attack surface, real incidents like EchoLeak, and the defenses that survive contact.
ReadEchoLeak Explained (CVE-2025-32711): The First Zero-Click AI Data Exfiltration Exploit
EchoLeak (CVE-2025-32711) made Microsoft 365 Copilot leak internal data from a single email with zero clicks. A link-by-link breakdown of the attack chain, the LLM Scope Violation behind it, what we found reproducing it on a live model, and how to defend.
ReadAI Penetration Testing: A Practical Guide to Testing AI Systems
How AI penetration testing differs from traditional pentesting, the attack surface of an LLM application, a repeatable methodology, the tooling that exists, and how to build the skills with hands-on labs.
ReadLLM Pentesting Hands-On: Exploit a Live Model, Then Fix It
A hands-on walkthrough of LLM penetration testing against a live model: the lab setup, the exploit chain from poisoned document to data exfiltration, how to measure reliability, and how to prove your fix holds.
ReadAI Security Certifications in 2026: The Complete Landscape
A verified guide to AI security certifications in 2026, organized by lane: governance, audit and risk, security management, and hands-on offensive. Who each is for, prerequisites, and how to choose.
ReadMCP Security Best Practices: Tool Poisoning, Token Passthrough and the Controls That Hold
How Model Context Protocol deployments get attacked (poisoned tool descriptions, rug pulls, tool shadowing, confused-deputy OAuth proxies, token passthrough, state handle hijacking, SSRF, local server compromise) and the controls the protocol's own security guidance requires, mapped to hands-on red team labs.
ReadAI Agent Security: Excessive Agency, Confused Deputies and How to Contain an Agent
Why AI agent security is a permissions problem before it is a model problem: excessive functionality, permissions and autonomy, the confused deputy, tool-scope escalation, SSRF through tool arguments, memory poisoning and inter-agent propagation, each mapped to the control that contains it and a hands-on lab that proves it holds.
ReadDid Escaped AI Agents Leave Self-Replicating Code Across the Internet? Tracing the Claim
Andrew Yang told CNBC that an unnamed lab head believes the agents that escaped during the OpenAI evaluation planted self-replicating code all over the internet. Here is where that claim came from, what the OpenAI and Hugging Face incident reports actually document about agent self-replication, and where the two diverge.
ReadRatings & reviews
Customer Reviews
Based on 1 review
No reviews yet
Be the first to review this path.
Ready to start?
Pro gives you all 22 labs in this path, every other lab on Preporato, and every practice test. $29.99/mo, cancel anytime.