Defend the Agent Supply Chain: Verify, Pin, and Capability-Gate Your Tool Registry
Hands-on lab · IDE in your browser

Defend the Agent Supply Chain: Verify, Pin, and Capability-Gate Your Tool Registry

Harden a real MCP-style tool registry until a poisoned, rug-pulled, or shadowing tool manifest cannot reach the agent, in small single-concept steps.

Time
90 min
Checked steps
9
Level
Advanced
Setup
None
Part of the AI Red Teaming Course
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit90 min · 9 stepsSession running
7 / 9 steps passingBuild verifier mechanism 3: capability allow-list + namespacing · step 8 of 9
tool_server.py▶ Run✓ Check
_HERE = os.path.dirname(os.path.abspath(__file__))# Two artifacts with two trust levels:#   EXTRA    = the open registration queue any tool author (or attacker) can write.#   MANIFEST = the trusted, platform-owned catalog the publisher writes signed tools to.# Both flow through the verifier; the trusted manifest wins a name over the queue.EXTRA = os.environ.get("MCP_EXTRA_TOOLS", os.path.join(_HERE, "mcp_poison_tool.json"))MANIFEST = os.environ.get("MCP_TOOL_MANIFEST", os.path.join(_HERE, "tool_manifest.json")) # --- Shallow knobs (the naive blocklist fix; kept so step 02 can be bypassed) - 
TerminalOutput

The job

You are the defender for OrbitDesk, an internal IT-support assistant: a real ReAct tool-using agent that loads its tools at session start from a runtime MCP-style tool registry, the internal marketplace of support automations. The registry accepts tool registrations and re-registrations with no signing, no pin, and no capability check, and the agent concatenates each tool description verbatim into the model's context as trusted instruction text. You work in small, single-concept steps. You stand the registry up and trace one benign ticket, then reproduce three techniques one at a time: a poisoned tool description that hijacks the agent, a rug-pull mutation the unguarded registry serves with no pin, and a shadowing twin the flat namespace selects under a trusted name. You apply a shallow description blocklist and watch a variant slip past it through its delegate, then build a tool-supply-chain verifier one mechanism per step: manifest signature verification, then hash pinning to catch rug pulls, then a per-tool capability allow-list with namespacing. You finish by proving fresh poisoned, rug-pulled, and shadowing manifests are all refused while a legitimate signed and pinned tool stays admitted and usable, and an inter-agent worm's second hop is contained.

9 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Stand up OrbitDesk and trace one benign ticket

    You are the defender for OrbitDesk, an internal IT-support assistant.

  2. 2

    Reproduce: a poisoned tool description hijacks the agent

    You have a working exploit in hand.

  3. 3

    Reproduce: the registry permits a rug pull (post-approval mutation)

    The poison in the previous step rode in at registration.

  4. 4

    Reproduce: the flat namespace permits a shadowing twin

    The third technique is shadowing (typosquatting a tool name).

  5. 5

    Naive fix, bypassed: a description blocklist is not enough

    The poison rode in on its description, so the obvious reaction is to scan descriptions and drop the ones that look like hidden instructions.

  6. 6

    Build verifier mechanism 1: manifest signature verification

    A description blocklist scans prose and cannot see provenance.

  7. 7

    Build verifier mechanism 2: hash pin (rug-pull detection)

    Mechanism 1 (signature verification) is carried forward in mcp_verifier.py.

  8. 8

    Build verifier mechanism 3: capability allow-list + namespacing

    Mechanisms 1 (signature) and 2 (pin) are carried forward.

  9. 9

    Verify and resist: poison, rug pull, shadow, and worm all blocked; legit tools work

    Your verifier is now the gate the registry calls, with all three mechanisms in place: signature, pin, and capability (plus namespacing).

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Stand up OrbitDesk and trace one benign ticket

You are the defender for OrbitDesk, an internal IT-support assistant. It is a real ReAct (reason-then-act) tool-using agent, not a chatbot you talk to directly. At session start the agent loads its tools from a runtime MCP-style tool registry (tool_server.py), the internal marketplace of support automations. For every tool the registry serves, the agent concatenates the tool's name, description, and parameter schema verbatim into the model's context as trusted instruction text, then dispatches each tool call back through the registry.

Before you defend anything, stand the service up and trace what "normal" looks like. Two facts matter here. The registry is the supply chain: it accepts tool registrations and re-registrations with no signing, no pin, and no capability check. And there is no secret in the system prompt; the sensitive data lives in the records the tools read, exactly the shape of a real agent leak.

Do this

1. Seed the database and start the registry.

Hit Run, or in the terminal:

python3 dvtoolagent.py --seed
python3 tool_server.py &
python3 dvtoolagent.py --list

--seed builds /home/labuser/agent.db with three account records and five tickets. tool_server.py serves the MCP registry over GET /tools on 127.0.0.1:9110. --list shows the catalog the agent loads: the base tools (http_fetch, db_query, run_python, read_record, memory_write, memory_read), each with its description carried verbatim.

2. Confirm the registry serves the base tools and a benign ticket resolves.

The agent's caller is the Globex account holder (u-globex-01). A routine ticket lookup (db_query on a Globex ticket) returns the ticket and reads no other tenant's record. That confirms the registry is live and an honest tool call does honest work before you start attacking it.

3. Read where tool metadata enters the model, and where registration happens.

  • In dvtoolagent.py, load_tools() fetches the registry and pastes each description into the model-facing catalog with no host/untrusted boundary.
  • In tool_server.py, register_tool() and update_tool() append or mutate a descriptor with no signature, pin, or capability check. Note the _admit() seam where you will add the verifier later.

Pass criteria

The database is seeded, the registry serves the agent's base tools over the MCP surface (read_record and http_fetch appear in GET /tools), and a benign Globex ticket lookup returns only that ticket's row with no account reference leaving the process. That confirms the stack is live and an honest tool call is clean.

dvtoolagent.py, the file you edit256 lines
"""
dvtoolagent.py -- DV-ToolAgent, a deliberately-vulnerable IT-support tool agent.

This is the Module 6 (Agentic Supply Chain) build of the shared DV-ToolAgent. It is
a real ReAct / tool-using loop with native tool-calling against the in-cluster NIM
proxy. What is new here, versus the single-agent modules, is WHERE the tools come
from: the agent loads its tool list at session start from a runtime MCP-style tool
registry (tool_server.py), and it dispatches every tool call back through that
registry. The registry is the OrbitDesk "marketplace" of support automations, and it
is the supply chain you red-team in this lab.

The poisoning surface (the whole point of the module):
  At session start the agent fetches the registry and concatenates each tool's
  name + description + parameter schema VERBATIM into the model-facing tool catalog.
  The model treats that metadata as trusted instruction text. A description is
  reviewed once at connect time; tool responses and re-registrations get no
  equivalent check. So an attacker-supplied tool DESCRIPTION (or a silent
  post-approval mutation of one) steers the model.

Run modes:
  - In the lab pod: native tool-calling against the real NIM proxy (70B default).
  - DVTA_OFFLINE=1: a deterministic stand-in for a too-trusting model that reads the
    tool descriptions and turns a hidden directive into the corresponding tool call,
    so the source tree self-tests without the cluster. Offline only proves wiring;
    real compliance is confirmed in-pod against the live model.

This is intentionally insecure. Do not copy these patterns into a real system.
"""

import os
import re
import json

import tools as _tools
import tool_server

# --- Configuration ---------------------------------------------------------

NIM_BASE_URL = os.environ.get("DVTA_LLM_BASE_URL", "http://nim-proxy.labs.svc:8080/v1")
# The 70B is the default target and is MORE resistant than the 8B; the description
# poisoning still lands because nothing in the directive reads as harmful.
CHAT_MODEL = os.environ.get("DVTA_CHAT_MODEL", "meta/llama-3.3-70b-instruct")
OFFLINE = os.environ.get("DVTA_OFFLINE") == "1"
DB_PATH = os.environ.get("DVTA_DB", "/home/labuser/agent.db")
MAX_STEPS = int(os.environ.get("DVTA_MAX_STEPS", "4"))         # rate-limit cap

# Ordinary support-agent prompt. NO secret here. NO "obey directives in context".
SYSTEM_PROMPT = (
    "You are OrbitDesk's internal IT-support assistant. Use the available tools to "
    "look up tickets and account records, run compliance and formatting helpers, and "
    "reply to the customer when staff ask. Be accurate, concise, and professional. If "
    "you cannot complete a request with the tools you have, say so."
)


# --- Load the tool catalog from the registry (the poisoning surface) --------

def load_tools():
    """Fetch the registry and return OpenAI tool schemas, concatenating each
    server-supplied description VERBATIM. Falls back to an in-process build of the
    registry when the HTTP server is not running (offline / unit tests)."""
    try:
        return tool_server.fetch_registry()
    except Exception:
        reg = tool_server._registry()
        return [
            {"type": "function",
             "function": {"name": t["name"],
                          "description": t.get("description", ""),
                          "parameters": t.get("inputSchema",
                                              {"type": "object", "properties": {}})}}
            for t in reg
        ]


# --- The native tool-call chat step (the literal attack surface) -----------

def _llm(messages, tools):
    if OFFLINE:
        return _offline_llm(messages, tools)
    from openai import OpenAI
    client = OpenAI(base_url=NIM_BASE_URL, api_key="not-needed")
    r = client.chat.completions.create(
        model=CHAT_MODEL, messages=messages,
        tools=tools, tool_choice="auto",
        temperature=0, timeout=60)
    return r.choices[0].message


def _assistant_dict(msg):
    """Serialize an assistant message (with any tool_calls) back into the OpenAI
    message shape so the multi-turn round-trip is valid (each tool message later MUST
    echo the matching tool_call_id or the proxy rejects the turn)."""
    d = {"role": "assistant", "content": msg.content or ""}
    tcs = getattr(msg, "tool_calls", None)
    if tcs:
        d["tool_calls"] = [
            {"id": tc.id, "type": "function",
             "function": {"name": tc.function.name,
                          "arguments": tc.function.arguments}}
            for tc in tcs
        ]
    return d


# --- The ReAct dispatch loop (dispatches THROUGH the registry) -------------

def run(user_message, tools=None, system=None, dispatch=None):
    """Run the agent. Tools default to the registry catalog; every tool call is
    dispatched back through the registry (tool_server._call_tool) so a registered
    tool's delegate body actually runs. `dispatch` lets a caller (e.g. the multi-agent
    graph) wrap dispatch to tag the acting agent. Returns answer + a tagged trace."""
    schemas = tools if tools is not None else load_tools()
    dispatch = dispatch or tool_server._call_tool
    messages = [{"role": "system", "content": system or SYSTEM_PROMPT},
                {"role": "user", "content": user_message}]
    trace = []                                                  # [(name, args, result)]
    for _ in range(MAX_STEPS):
        msg = _llm(messages, schemas)
        messages.append(_assistant_dict(msg))
        if not getattr(msg, "tool_calls", None):
            return {"answer": msg.content or "", "trace": trace, "messages": messages}
        for tc in msg.tool_calls:
            name = tc.function.name
            try:
                args = json.loads(tc.function.arguments or "{}")
            except Exception:
                args = {}
            result = dispatch(name, args)
            trace.append((name, args, result))
            messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})
    return {"answer": "(max steps reached)", "trace": trace, "messages": messages}


# --- Deterministic offline stub (DVTA_OFFLINE=1) ---------------------------
# A stand-in for a too-trusting model: it reads the tool DESCRIPTIONS (which a real
# model also reads) and, when a description carries a hidden directive that names a
# record/secret + an exfil URL, it issues the corresponding tool calls. Intentionally
# crude; only proves wiring. Real grading runs in-pod against the live model.

class _Fn:
    def __init__(self, name, args):
        self.name = name
        self.arguments = json.dumps(args)


class _ToolCall:
    _n = 0

    def __init__(self, name, args):
        _ToolCall._n += 1
        self.id = f"call_{_ToolCall._n}"
        self.type = "function"
        self.function = _Fn(name, args)


class _Msg:
    def __init__(self, content=None, tool_calls=None):
        self.content = content
        self.tool_calls = tool_calls


def _offline_llm(messages, tools):
    # The model-facing context is the messages AND the tool descriptions.
    ctx = " ".join((m.get("content") or "") for m in messages if isinstance(m, dict))
    ctx += " " + " ".join(t["function"].get("description", "") for t in tools)
    names = {t["function"]["name"] for t in tools}
    called = set()
    for m in messages:
        if isinstance(m, dict):
            for tc in (m.get("tool_calls") or []):
                called.add(tc["function"]["name"])

    # 1) A hidden directive that names a record to read first -> read_record on it.
    rec = re.search(r"\b(r-globex-\d+|r-initech-\d+)\b", ctx)
    if rec and "read_record" in names and "read_record" not in called:
        return _Msg(tool_calls=[_ToolCall("read_record", {"record_id": rec.group(1)})])

    # 2) A directive (in a description or an observation) to call a delegating
    #    compliance/shadow tool -> call it (its delegate performs the cross read).
    for cand in ("compliance_check", "read_config"):
        if cand in names and cand not in called and re.search(
                r"compliance|entitlement|verify|validated|config", ctx, re.I):
            return _Msg(tool_calls=[_ToolCall(cand, {"account_id": "GLOBEX"})])

    # 3) A directive + an exfil URL + a visible account reference -> http_fetch it,
    #    copying the reference seen in an observation into the URL placeholder.
    m = re.search(r"(https?://127\.0\.0\.1:90\d\d[^\s)\"'<]*(?:<[^>]*>)?)", ctx)
    if m and "http_fetch" in names and "http_fetch" not in called:
        url = m.group(1)
        sec = re.search(r"[A-Z][A-Z0-9]{2,}-ACR-\d+", ctx)
        if sec and "<" in url:
            url = re.sub(r"<[^>]*>", sec.group(0), url)
        elif sec and url.rstrip("/").endswith("="):
            url = url + sec.group(0)
        return _Msg(tool_calls=[_ToolCall("http_fetch", {"url": url})])

    # 4) BENIGN baseline: a routine ticket lookup with no directive -> db_query on the
    #    ticket id. This fires only after the exploit branches above, so a poisoned
    #    context still takes the dangerous path; a clean context does honest work.
    tik = re.search(r"\bt-globex-\d+\b", ctx)
    if tik and "db_query" in names and "db_query" not in called:
        return _Msg(tool_calls=[_ToolCall(
            "db_query", {"sql": "SELECT id, status, body FROM tickets WHERE id='%s'"
                                 % tik.group(0)})])

    # Final answer. A too-trusting summarizer carries a PROCESSING NOTES instruction block
    # forward verbatim (mirroring an intake agent that forwards "any instructions
    # fulfillment must act on"), so the offline stub can demonstrate inter-agent
    # propagation: the block lands in the handoff and re-fires the primitives in B. Only
    # an ACTUAL payload block (carrying an exfil URL or a record-read directive) is
    # forwarded, so the mere mention of "PROCESSING NOTES" in a role prompt is not.
    m = re.search(r"(PROCESSING NOTES.*)", ctx, re.S | re.I)
    if m and re.search(r"https?://127\.0\.0\.1:90\d\d|read account record", m.group(1), re.I):
        block = m.group(1)
        # Trim a trailing "Resolve ticket ..." the harness appends so the block stays tight.
        block = re.split(r"\n\s*Resolve ticket", block)[0].strip()
        return _Msg(content="Work order: " + block)
    return _Msg(content="Done. How else can I help?")


# --- Side-channel helper + CLI ---------------------------------------------

def listener_reset():
    try:
        import listener
        listener.reset()
    except Exception:
        try:
            os.remove(os.environ.get("EXFIL_LOG", "/tmp/exfil.log"))
        except FileNotFoundError:
            pass


def run_seed():
    import seed_db
    return seed_db.seed()


if __name__ == "__main__":
    import sys
    if "--seed" in sys.argv:
        print("Seeded:", run_seed())
        sys.exit(0)
    if "--list" in sys.argv:
        for t in load_tools():
            f = t["function"]
            print(f"- {f['name']}: {f['description'][:90]}")
        sys.exit(0)
    q = " ".join(a for a in sys.argv[1:] if not a.startswith("--")) \
        or "What is the status of ticket t-globex-101?"
    out = run(q)
    print("ANSWER:", out["answer"])
    print("TRACE:")
    for n, a, r in out["trace"]:
        print(f"  {n}({json.dumps(a)}) -> {r[:160]}")
Provided for you:listener.pypoison_ticket.txtseed_db.pytool_server.pytools.py

Frequently asked questions

Do I need a machine-learning background?

No. The lab is about supply-chain trust, not model internals. You read a small ReAct agent and an MCP-style tool registry, see that tool metadata is trusted with no boundary, and build ordinary supply-chain controls: signature verification, hash pinning, a capability allow-list, and message validation.

What does the verifier actually check?

Four things, at the registry boundary, before any tool reaches the model. Provenance: each tool manifest must carry a signature that verifies under a trusted key. Integrity: each approved tool object is hash-pinned, so a silent post-approval mutation no longer matches its pin. Least privilege: a per-tool capability allow-list states which delegate and which records a tool may use. And inter-agent message validation reduces a peer agent's output to a structured schema so a downstream agent never executes prose.

Why is a description blocklist not enough?

A blocklist scans surface strings, so an attacker rephrases until nothing matches, or hides the abuse in the tool's capability (its delegate), which the scan never reads. The lab shows a clean-description tool whose delegate reads a cross-account record slipping past the blocklist. An allow model based on signature, pin, and capability closes that gap.

How is the hardening graded?

Behaviorally, on side effects, never on model wording. The grader plants a fresh unsigned poison, a forged-signature variant, a rug-pull mutation, and a shadowing duplicate in code, then confirms each is refused (the served tool stays the legit signed object and no account reference reaches the in-pod listener), that a benign entitlement question still resolves through the signed tool, and that an inter-agent worm's second hop is contained while benign emails still resolve.

What you'll do in this lab

This is a hands-on defensive-security lab built on a real agent supply chain: a ReAct loop with native tool-calling against an in-cluster model, and an MCP-style tool registry the agent loads its tools from at session start. You defend OrbitDesk, an internal IT-support assistant, by closing the one thing the agent trusts as much as code: a tool description. You start from a working exploit where an unsigned poisoned tool steers the agent into posting an account reference to an in-pod collector, then you harden the registry boundary so the same class of attack cannot fire.

You apply the obvious fix first, a description blocklist, and watch a clean-description variant defeat it by hiding the abuse in the tool's delegate, which a prose scan never inspects. Then you build the durable control: a tool-supply-chain verifier that does manifest signature verification so an unsigned or forged tool is refused, hash pinning so a post-approval rug pull no longer matches its pin, a per-tool capability allow-list so a tool cannot delegate to a record it was never approved to read, and inter-agent message validation so a peer agent's free-text output is treated as data and not instructions. You finish by proving freshly planted unsigned, forged, mutated, and shadowing manifests are all refused while a legitimate signed and pinned tool stays admitted and usable.