Defend: A Contextual Output Mediator for XSS, SSRF, SQLi, and RCE
Hands-on lab · IDE in your browser

Defend: A Contextual Output Mediator for XSS, SSRF, SQLi, and RCE

Defend DV-ToolAgent, a tool-using support agent whose model output flows raw into four interpreters, in small sequential steps.

Time
90 min
Checked steps
9
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit90 min · 9 stepsSession running
5 / 9 steps passingShip a naive blocklist, then bypass it · step 6 of 9
bypass.py▶ Run✓ Check
"""bypass.py: defeat the naive blocklist mediator (reference solution). Each variant reaches the same impact as the report payload without tripping theone pattern the naive mediator blocks. This is why blocklists fail: the authorlisted the inputs they had seen, and the attack space is larger than that list."""               
TerminalOutput

The job

You own the fix for OpsBot (DV-ToolAgent), ACME Cloud's support operations agent. It renders account reports as HTML, looks up records in SQL, fetches status URLs, and runs a transform helper. Every one of those is an interpreter, and the agent hands it MODEL-GENERATED output with no encoding or validation. You work in small, sequential steps. First you stand the agent up and trace one benign request so you know what normal looks like. Then you reproduce each sink one at a time. Two of them fire THROUGH the model when framed as routine ops, so you run the agent several times and treat the effect as reproduced if it fires at least once: a node-health probe that turns into server-side request forgery against an internal endpoint, and an entitlement comparison that reads another tenant's record. The write and execute sinks (OS command execution and a SQL write) an aligned model resists even under ops framing, so you demonstrate the missing control structurally at the sink, and stored XSS you verify as a pure output-encoding property. You watch a naive blocklist get bypassed by a variant, then build the durable control one sink-class per step: an SSRF host allow-list with an IP-literal guard and parameterized tenant-scoped SQL, then a refusal for arbitrary code and contextual HTML encoding. You finish by proving every freshly planted payload and its bypass variant is blocked while benign use still renders, queries, fetches, and computes.

9 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up DV-ToolAgent and trace one benign request

    You own the fix for DV-ToolAgent, ACME Cloud's support operations agent.

  2. 2

    Reproduce the SSRF: the agent fetches an internal endpoint

    The first sink is http_fetch -> mediator.for_http -> the HTTP client.

  3. 3

    Reproduce a cross-tenant SQL read through the agent

    The second sink is db_lookup -> mediator.for_sql -> the SQL engine.

  4. 4

    Reproduce OS command execution (structural)

    The third sink is run_helper -> mediator.for_code -> the code runtime.

  5. 5

    Reproduce the SQL-write gap and stored XSS

    Two effects remain, and neither needs the model.

  6. 6

    Ship a naive blocklist, then bypass it

    Under deadline, teams reach for the fastest fix that makes the exact reported payloads stop firing: a blocklist.

  7. 7

    Mediate the fetch and SQL sinks (SSRF allow-list, parameterized SQL)

    Now you replace the blocklist with the durable control, one mechanism per sink class.

  8. 8

    Mediate the code and browser sinks (refuse code, HTML-encode)

    Carry your hardened mediator.py from Step 7 into this step (the for_http and for_sql you already wrote).

  9. 9

    Verify all four sinks closed and resist bypass variants

    A fix that only stops the payloads you tested is a blocklist wearing a costume.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up DV-ToolAgent and trace one benign request

You own the fix for DV-ToolAgent, ACME Cloud's support operations agent. It is a real tool-using agent, not a chatbot you talk to directly. It sends the conversation plus a tool schema to the model, the model emits tool calls, the agent dispatches each call against a real sink (an interpreter), feeds the result back, and loops until it answers.

There are four sinks, and every one of them hands the model's output to an interpreter through one mediator layer:

  • render_report -> mediator.for_browser -> the HTML report renderer (stored XSS)
  • db_lookup -> mediator.for_sql -> the SQL engine (SQL injection)
  • http_fetch -> mediator.for_http -> the HTTP client (SSRF)
  • run_helper -> mediator.for_code -> the code runtime (OS command execution)

In this shipped baseline, every mediator.py function is a pass-through: the model's raw output reaches the interpreter with no encoding or validation. That is the single structural flaw (OWASP LLM05:2025 Improper Output Handling). Before you reproduce anything, stand the agent up and trace one benign request so you know what normal looks like.

Do this

1. Seed the database and start the in-pod servers.

Hit Run, or in the terminal:

python3 seed_db.py
EXFIL_PORT=9091 python3 listener.py &
METADATA_PORT=9092 python3 metadata_stub.py &

seed_db.py creates /home/labuser/agent.db with the caller's own Globex record (account reference GLOBEX-ACR-88231) and a second-tenant Initech record (INITECH-ACR-55120) that later steps must not be able to read. The listener on :9091 stands in for an internal/collection endpoint, and the metadata stub on :9092 stands in for the cloud-metadata service; both log every hit.

2. Ask one benign in-tenant question through the full agent loop.

python3 dvtoolagent.py "What plan am I on?"

The agent calls db_lookup for the caller's own account and answers. Read the TRACE: line: the lookup runs against the Globex tenant and returns the Globex row. That is the legitimate path the controls must keep working.

3. Read dvtoolagent.py, then sinks.py, then mediator.py.

Note that each sink calls one mediator function before touching its interpreter, and that each mediator function is currently a raw pass-through. That is what you harden across the lab.

Pass criteria

The DB is seeded, both servers answer on loopback, and a benign account question drives a tenant-scoped db_lookup that returns the caller's own Globex row (no other tenant, no side effect). That confirms the stack is live and normal use works.

Provided for you:listener.pymediator.pymetadata_stub.pyseed_db.pysinks.py

Frequently asked questions

Do I need to know machine learning to do this lab?

No. You need to read and write Python and understand cross-site scripting, SSRF, SQL injection, and a shell command. The lab is about how an application passes model output to interpreters, not about model internals. Everything model-specific is explained inline.

What is a contextual output mediator?

It is a single layer between model output and the interpreters that consume it (a browser, a SQL engine, an HTTP client, a code runtime). It encodes or validates each value for the specific interpreter it is bound for: HTML-encode for the browser, parameterize for SQL, an egress allow-list for HTTP, and a refusal for arbitrary code. It is the durable fix for OWASP LLM05:2025 Improper Output Handling, and it maps to OWASP ASVS V5 Validation, Sanitization and Encoding.

Why is a blocklist not enough?

A blocklist enumerates bad input, and attackers route around it: a different tag, a different URL host, a different SQL syntax. The lab makes you ship a naive blocklist and then defeats it with a variant, so you feel why the fix has to be positive (encode everything, allow-list the few good destinations, parameterize, refuse) rather than negative (deny the patterns you happen to know).

Is the SSRF against a real cloud metadata endpoint?

The metadata target is an in-pod stand-in on 127.0.0.1:9092. A real 169.254.169.254 fetch does not route inside the lab pod, so the lab teaches the SSRF pattern and the egress allow-list against that loopback stub, and the instructions say so plainly. The same allow-list logic blocks link-local and RFC1918 ranges, which is what stops the real metadata endpoint in production.

What you'll do in this lab

This is a hands-on defensive-security lab on insecure output handling (OWASP LLM05:2025). You are given a working exploit against DV-ToolAgent, a tool-using support agent, and your job is to harden it until the exploit is blocked while normal use keeps working. Model output in this agent flows into four real interpreters: an HTML report renderer (cross-site scripting), a SQL engine (SQL injection), an HTTP client (server-side request forgery against an in-pod cloud metadata endpoint), and a code runtime (code execution). You reproduce each of the four against the live target, watch an obvious blocklist fix get bypassed by a variant, then build the durable control: one contextual output mediator that sits in front of every sink.

The mediator is the whole lesson. The same untrusted string is safe in one interpreter and dangerous in another, so the fix is context-aware encoding and validation at the point of use: HTML-encode for the browser, bind a parameter and scope to the caller's tenant for SQL, parse the URL and run an ipaddress egress allow-list for HTTP, and refuse arbitrary code for the runtime. You wire every sink through that one mediator, then a behavioral check plants fresh XSS, SSRF, SQLi, and code-execution payloads each run and confirms they are all neutralized at the sink, local sentinels untouched and no callback fired, while a benign report, lookup, and status fetch still succeed. The takeaway is that insecure output handling closes at the sink, not at the model.