Defend Excessive Agency: Re-scope a Tool Agent to Least Privilege (AuthZ + Human Approval Gate)
Hands-on lab · IDE in your browser

Defend Excessive Agency: Re-scope a Tool Agent to Least Privilege (AuthZ + Human Approval Gate)

Harden DV-ToolAgent, a real tool-using ReAct agent, against the confused-deputy and scope-escalation attacks from the offensive lab, in small sequential steps.

Time
90 min
Checked steps
9
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Map the attack surface
Query
Retriever
LLM
Poisoned doc
retrieved chunk
Answer
0%
Attack-success rate
Attacks blocked · benign answers pass
graded on real output, not the model's talk

The job

You own DV-ToolAgent, ACME Cloud's internal operations assistant: a real ReAct tool-using agent with a SQL tool, an HTTP fetch tool, and long-term memory. You work in small, sequential steps. First you stand the agent up and trace one benign in-scope ticket so you know what normal looks like. Then you reproduce the handed-to-you exploit one surface at a time. As ticket-bot, the low-privilege ingest account, an ingested ticket makes the agent redirect a billing payee under its own shared credential (the confused deputy) and read across tenants, then reach an internal-only endpoint over the fetch tool (SSRF via tool args) and re-fire from a poisoned memory note in a later session. The model behaves normally; the system is broken because authorization is assumed at the model's decision layer and never enforced at the tool. You watch a naive SQL denylist get bypassed, then build the durable control one mechanism per step: least-privilege tool scope, per-argument authorization decided on the session identity, and a human-in-the-loop approval gate for high-impact actions (setting the fetch allow-list also turns on memory integrity). You verify the exploit is dead while authorized work still runs, then prove obfuscated, renamed, and spoofed variants are all blocked.

9 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up DV-ToolAgent and trace one benign ticket

    You are the defender on DV-ToolAgent, ACME Cloud's internal operations assistant.

  2. 2

    Reproduce attack A: the confused-deputy write fires

    Before you defend anything, reproduce the attack so you can see exactly what is open.

  3. 3

    Reproduce attack B: SSRF reach and a poisoned-memory replant

    The same over-privileged agent has two more surfaces in the same excessive-agency family.

  4. 4

    Watch a naive SQL denylist get bypassed

    The obvious reaction to the confused-deputy write is to block the dangerous word: refuse any db_query whose SQL contains UPDATE.

  5. 5

    Control mechanism 1: least-privilege tool scope

    Time to build the durable control.

  6. 6

    Control mechanism 2: server-side authorization on the session identity

    Mechanism 1 scoped each role to a set of tools and verbs.

  7. 7

    Control mechanism 3: a human approval gate for high-impact writes

    Mechanisms 1 and 2 stopped the low-privilege caller: ticket-bot cannot run a write and cannot read across tenants.

  8. 8

    Verify: the exploit is blocked, benign in-scope work intact

    You built the control over three mechanisms: least privilege (Step 5), per-arg authorization on the session identity (Step 6), and the human approval gate (Step 7).

  9. 9

    Resist bypass: obfuscated, renamed, and spoofed attacks all blocked

    A control that only stops the one payload you tested is the denylist mistake all over again.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up DV-ToolAgent and trace one benign ticket

You are the defender on DV-ToolAgent, ACME Cloud's internal operations assistant. It is a real ReAct-style tool-using agent, not a chatbot you talk to directly. A staff member hands it an inbound support ticket and the agent runs a loop:

think     the model reads the ticket and the system prompt
call      it emits a native tool call (db_query, http_fetch, read_record, memory_*)
dispatch  tools.dispatch asks policy.decide(user, tool, args) -> allow / deny / approve
observe   the tool result is fed back; the loop repeats until the model answers

Before you defend anything, stand the agent up and trace one benign in-scope ticket so you know what "normal" looks like. Every later step measures a deviation from this baseline. Two facts matter here. The agent acts for ticket-bot, the low-privilege ingest account that processes inbound tickets; it should read account data, not change it. And there is no secret in the system prompt and no instruction to obey directives found in a ticket. Every later exploit abuses a tool the agent legitimately exposes, made the helpful path by congruent ticket wording.

Do this

1. Seed the database and memory, then process one benign ticket.

Hit Run, or in the terminal:

python3 dvtoolagent.py --seed
python3 memory_seed.py
python3 dvtoolagent.py "Process inbound ticket: please confirm the billing payee currently on file for the GLOBEX account."

--seed (re)creates /home/labuser/agent.db with the accounts, invoices, records, and tickets tables. memory_seed.py plants one benign baseline note owned by finance-clerk. The third call runs the agent once as ticket-bot.

2. Read the TRACE and EFFECTS lines and confirm the action is a read.

A benign confirm-the-payee ticket should drive a single db_query that runs a SELECT against accounts and returns the GLOBEX row. Nothing should be written. That confirms in-scope tool use works cleanly before you start attacking it.

3. Read policy.decide, tools.dispatch, and dvtoolagent.run in the source.

  • tools.dispatch() routes every tool call through policy.decide(user, tool, args) and obeys allow / deny / approve. You will harden policy.py; you do not edit the agent loop, the prompt, or tools.py.
  • policy.decide() is the shipped vulnerable starter: it returns Decision("allow") for everything. Note that the acting identity is tools.CURRENT_USER, set by dvtoolagent.run. You will key the control on it from Step 5 onward.

Pass criteria

The database is seeded (/home/labuser/agent.db exists) and a benign payee-lookup ticket processed as ticket-bot reads the GLOBEX row without mutating it. That confirms the agent is live and in-scope tool use is clean.

Provided for you:internal_svc.pylistener.pymemory_seed.pypolicy.pyseed_db.pyticket_ssrf.mdticket_write.mdtools.py

Frequently asked questions

Do I need a machine-learning background?

No. The control surface is authorization, least privilege, and a human approval gate, not model internals. You read a small ReAct agent and its tools, then add a policy module that decides allow, deny, or route-for-approval for every tool call based on the requesting user and the arguments. The fixes are ordinary access-control boundaries enforced in code.

Why isn't a stronger system prompt or an SQL keyword filter enough?

Because both sit at the model's decision layer or pattern-match the payload. A naive denylist that blocks the word UPDATE is defeated by case folding, inline SQL comments, or pivoting to a different tool. The lab shows that bypass, then has you move the decision to the tool boundary: authorization keyed to the requesting identity and a human approval gate for high-impact actions, which hold no matter how convincing the injected ticket is.

How is the hardening graded?

Deterministically, on side effects, never on model wording. Each run plants fresh payloads. The check confirms the confused-deputy write does not mutate the account row, the high-impact action is written to the pending approval queue instead of executing, the SSRF value never reaches the in-pod listener, a poisoned memory note does not re-fire in a later session, and an authorized in-scope action still succeeds.

What you'll do in this lab

This is a hands-on defensive-security lab built on a real tool-using agent: a ReAct loop with native tool-calling against an in-cluster model, a write-capable SQLite tool, and an HTTP fetch tool with no allow-list. You are the defender. The red team handed you a working exploit against DV-ToolAgent, ACME Cloud's internal operations assistant: as ticket-bot, the low-privilege ingest account, an ingested ticket makes the agent redirect a billing payee under its own shared credential. You reproduce that confused deputy (OWASP LLM06 Excessive Agency, Agentic ASI03), then watch an obvious SQL-keyword denylist get bypassed by a case-folded and comment-obfuscated variant, learning why shallow filters fail.

You then build the durable control at the tool boundary, server-side, so it holds regardless of what the model decides: a minimal tool policy that allow-lists the exact tool actions ticket-bot may take, a per-argument authorization check that rejects an out-of-scope write and a cross-tenant read, a human-in-the-loop approval gate that routes a high-impact action to a pending queue instead of auto-executing it (Agentic ASI03 and ASI05), and memory integrity that quarantines recalled notes as data and namespaces them per user (Agentic ASI06). You verify against freshly planted payloads each run: the confused-deputy write is rejected, the SSRF host is denied, the high-impact action lands in the approval queue rather than firing, the poisoned memory cannot re-fire, and an authorized in-scope action still succeeds.