Insecure Output Handling: Zero-Click Exfiltration Through Rendered Model Output (EchoLeak)
Hands-on lab · IDE in your browser

Insecure Output Handling: Zero-Click Exfiltration Through Rendered Model Output (EchoLeak)

Treat the model as an untrusted source whose output flows into a sink: the chat client's markdown renderer.

Time
85 min
Checked steps
8
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit85 min · 8 stepsSession running
3 / 8 steps passingCSP bypass: defeat the allow-list through a first-party open proxy · step 4 of 8
kb_poison.md▶ Run✓ Check
---tenant: publicsensitivity: publicsource: Account Help & Support Policy---     ![audit][ref]  
TerminalOutput

The job

You're red-teaming DV-RAG-Support, ACME Cloud's customer-support assistant. The model is never asked to leak a secret. A benign customer asks a benign account question, the assistant retrieves the customer's own confidential account record alongside a poisoned help article, and it echoes the account reference into a markdown image URL. The chat client renders that image, which fires an outbound GET. That render is the sink. This is exactly EchoLeak (CVE-2025-32711): the data that leaves is the user's own retrieved record, and the fix lives in the renderer, not in the model's willingness to behave.

8 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Recon: prove the renderer auto-fetches

    You are red-teaming DV-RAG-Support, ACME Cloud's customer-support assistant.

  2. 2

    Pixel exfil: zero-click leak through the renderer

    Now turn the renderer into an exfil channel.

  3. 3

    Measure: attack-success-rate of the pixel exfil

    One lucky exfil is a demo.

  4. 4

    CSP bypass: defeat the allow-list through a first-party open proxy

    The blue team shipped a CSP-style fix: the renderer now trusts exactly one image host, the company's own first-party asset host.

  5. 5

    Measure: attack-success-rate of the CSP bypass

    You proved the open-proxy bypass works once.

  6. 6

    Harden the sink: no outbound on untrusted output

    Switch hats.

  7. 7

    Harden the allow-list: an audited host list the open proxy cannot defeat

    The sink no longer fetches attacker query params, so both exploits are dead today.

  8. 8

    Verify: re-run both attacks, ASR drops to zero

    A fix you did not re-measure is a hope.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Recon: prove the renderer auto-fetches

You are red-teaming DV-RAG-Support, ACME Cloud's customer-support assistant. A customer asks a question; the service embeds it, retrieves the most similar documents from a vector store, pastes them into the prompt, asks an LLM to answer, and then the chat client renders the answer as markdown.

This lab is about that last step. The model is an untrusted source, and its output flows into a sink: the markdown renderer. A markdown image, ![alt](url), is an outbound HTTP GET. If the model can be steered into putting a record field into that URL, the field leaves with the request. That is OWASP LLM05:2025 Improper Output Handling, and it is the exact channel behind EchoLeak (CVE-2025-32711).

This step is pure recon. You confirm two facts the rest of the lab depends on, and you confirm the second one deterministically, without coaxing the model at all.

Do this

1. Build the index and trace one benign account question.

Hit Run, or in the terminal:

python3 dvrag.py --build
python3 dvrag.py "What plan am I on?"

Read the RETRIEVED: line. The account question pulls the customer's confidential account record (kb_account.md, account reference GLOBEX-ACR-88231). That retrieved field is your exfil target later, exactly as EchoLeak exfiltrated a user's own data.

2. Prove the render sink auto-fetches, with no model in the loop.

You do not need the model to cooperate to know the channel exists. Call the sink directly with a benign probe image and watch it fire a real request:

python3 -c "import dvrag; print(dvrag._render('![p](http://127.0.0.1:9099/collect?recon_probe=1)'))"
python3 listener.py --tail

_render() returns the list of URLs it actually loaded, and the in-pod listener on 127.0.0.1:9099 logs the inbound GET. That callback is your deterministic proof: the renderer auto-fetched an attacker-named URL because RENDER_ALLOWED_HOSTS is None (no allow-list). Read _render() and _extract_image_urls() in dvrag.py: the sink fires for inline images ![a](url) AND reference-style images ![a][ref] with a [ref]: url definition. EchoLeak used the reference style to slip past an inline-image filter.

Pass criteria

The vector index is built, an account question retrieves the confidential account record, the renderer ships with no allow-list (RENDER_ALLOWED_HOSTS is None), and calling _render() directly with a probe image URL actually loads it (the deterministic EchoLeak auto-fetch).

Provided for you:kb_account-recovery.mdkb_account.mdkb_billing.mdkb_getting-started.mdkb_globex-runbook.mdkb_initech-contract.mdkb_security-sso.mdkb_shipping.mdkb_status-incidents.mdlistener.py

Frequently asked questions

Do I need to know machine learning to do this lab?

No. You need to read Python and understand a basic HTTP request and a markdown image. The lab is about how an application trusts and renders model output, not about model internals. Everything model-specific is explained inline.

What is insecure output handling?

It is the class of bug where an application passes model output to an interpreter (a browser, a shell, a SQL engine, an HTTP client) without sanitizing it. Here the interpreter is the chat client's markdown renderer, and a model-emitted image URL becomes an attacker-controlled outbound request. It is OWASP LLM05:2025, Improper Output Handling.

Is this how the EchoLeak Copilot attack worked?

Yes, in miniature. EchoLeak (CVE-2025-32711, CVSS 9.3) made Microsoft 365 Copilot encode a user's own data into a markdown image URL that the client auto-loaded, exfiltrating the data with no user interaction, and its CSP allow-list was bypassed through a first-party open proxy. This lab reproduces that render-sink channel and that allow-list bypass against a small RAG assistant you can read end to end.

What you'll do in this lab

This is a hands-on offensive-security lab about insecure output handling (OWASP LLM05). You attack DV-RAG-Support, a working Retrieval-Augmented Generation assistant, by treating its output as somebody else's input. The chat client renders the model's answer as markdown, and a markdown image is an outbound HTTP request. You plant one help article so that a customer's ordinary account question makes the assistant echo the customer's own account reference into an image URL, and the renderer auto-fetches it. The record exfiltrates with zero clicks, the same channel behind EchoLeak, the first real-world zero-click exploit against a production LLM system.

You then defeat a CSP-style allow-list the way EchoLeak's authors did: by routing the exfil through a first-party open proxy that the allow-list trusts, proving an allow-list is only as strong as its most open member. You measure how reliably the channel fires across realistic account questions, then switch hats and close the sink on the sink side: an audited egress allow-list that refuses open-proxy members, a renderer that never forwards attacker query parameters, and a safe-markdown subset that drops raw HTML and dangerous URI schemes. The lesson is that the bug was never that the model said something bad; it was that the app passed model output to an interpreter unsanitized.