Track · AI agents and MCP
Agentic AI course: build AI agent projects hands-on
Agents from the tool loop up: write the loop by hand, then ReAct, MCP servers and clients, long-term memory, supervisors of specialist agents and human approval steps, and finish by evaluating every trajectory and routing between models.
- 26
- Labs
- 22 h
- In total
- Beginner to advanced
- Level
- 2
- Free
What you will build
- An agent tool loop written by hand, then the same agent on ReAct with NVIDIA NIM
- An MCP tool server and an MCP client, with discovery, errors and approvals handled
- Agents with long-term memory, a supervisor of specialists and a human approval step
- Agents that research the web, query a database and analyse data with pandas, safely
- An evaluation harness that scores trajectories and tool calls, and a gate in front of risky tools
Before you start
- Python basics; you read and edit short scripts
- What a tool call is: the model emitting a function name plus arguments
- No prior LangGraph experience; the labs introduce nodes, edges and state as they go
Tools you will use
LangGraphLangChainNVIDIA NIMMCPA2AMilvusDSPypandasSQL
Labs in this track
In order, from the first lab to the hardest. Every lab stands on its own, so start wherever you like.
One agent, done right
Write the tool loop without a framework (free), build a ReAct agent, compare patterns, then point agents at the web, pandas and a database.
- Lab 1Build an Agent From Scratch: The Tool Loop Without a FrameworkHand-roll a tool-using agent on the chat completions API and nothing else: parse the model's tool calls, dispatch them to plain Python functions, close the loop with correctly threaded tool results, add stop conditions and budgets, turn every failure into a tool result the model can read, put an approval guardrail in front of a side-effecting refund tool, control context with truncation and a trace, and finish with an evaluation run that scores the agent on a fixed question set.80 minIntermediateHostedFree# agent-tool-loop-from-scratch · agentPOST /api/agent/invoke200 OK · graded
- Lab 2Build a ReAct Agent with NVIDIA NIMBuild an AI research librarian: an agent that searches a corpus of ML papers, compares methods and answers multi-step questions on real NIM endpoints with LangGraph.35 minIntermediateHostedPro# ReAct · thought/action/observe1. Thought:2. search_docs(…)3. Observation:Final answer → 200 OK
- Lab 3Build an AI Agent 3 Ways: ReAct vs Tool Calling vs Plan-and-ExecuteBuild the same SaaS support agent three ways, as ReAct, direct tool calling and plan-and-execute, then compare speed, reasoning quality and reliability to learn when each fits.35 minIntermediateHostedPro
- Lab 4Structured Output & Function Calling with NIMGet reliable machine-parseable data out of an LLM. Compare prompt-only JSON extraction with the function-calling API, chain two tools and measure the reliability gap.30 minIntermediateHostedPro# structured-output-tools · agentPOST /api/agent/invoke200 OK · graded
- Lab 5Web Research Agent: Search, Read and Cite Without Being Fooled by the WebBuild a research agent that searches the web, reads pages and answers with citations it can prove. Clean pages into text, run a tool-calling loop with a step limit, verify every citation against the pages actually read, hold off prompt injection planted in a page, and prefer official, current sources over spam, stale blogs and forum posts.50 minIntermediateHostedPro# web-research-agent · agentPOST /api/agent/invoke200 OK · graded
- Lab 6Code-Executing Data Agent: Let an LLM Answer Questions With pandas, SafelyBuild an agent that answers business questions about a CSV by writing pandas code and running it. Run the code in a sandboxed process with time, memory and output limits, describe the data so the model writes working code, insist that every answer is a number the code printed, refuse dangerous code with an AST allow-list, and move business rules out of the prompt into code.50 minIntermediateHostedPro# data-agent · agentPOST /api/agent/invoke200 OK · graded
- Lab 7Text-to-SQL Agent: Answer Questions From a Database Without Letting the Model Touch the DataBuild a text-to-SQL assistant on SQLite that is safe to point at real data. Describe the schema with its business-rule comments, run queries on a read-only connection with an authorizer, time limit and row cap, reject anything but one compiling SELECT and let the model repair its errors, grade answers by their result rows, and add retrieved example queries that teach the rules.50 minIntermediateHostedPro# text-to-sql-agent · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 8Spec to Code: Build a Test-Driven Coding AgentBuild the loop a coding agent runs: turn a specification into tests, generate a solution, run the tests in a sandbox, feed the failures back and try again until they pass, then check the result against held-out tests so a green run means the problem was solved and not gamed. The model is scripted in the checks and real when you run it.55 minIntermediateHostedPro# spec-to-code · agentPOST /api/agent/invoke200 OK · graded
- Lab 9AI Code Review Bot: Parse a Diff, Find Bugs, Score and Gate the PRBuild a code review bot around a model. Parse a pull-request diff into the lines that changed, prompt the model for structured findings and read them back through prose and code fences, score what it caught against the bugs planted in the fixtures, triage the noise into a ranked list, and turn it into a block-or-approve verdict with a review comment.50 minIntermediateHostedPro# review-bot · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 10Refactor Legacy Code Safely: Characterization Tests and an AI Refactoring AgentBuild an agent that refactors messy legacy code without changing its behaviour: pin the current behaviour with a characterization suite, build a regression net that runs any candidate against the golden results, prompt a model to clean the code up while preserving every quirk, wrap it in a loop that feeds regressions back and fails safe, and enforce the rule that a refactor which changes a single output is rejected even if it looks more correct.55 minIntermediateHostedPro# refactor-agent · agentPOST /api/agent/invoke200 OK · graded
- Lab 11Debug With an Agent: Reproduce, Fix, and Add the Regression TestBuild an agent that debugs a failing function the disciplined way: reproduce the failure before touching the code, prompt a model for a fix from the buggy source and the failing tests, verify the fix against the whole suite, loop while feeding failures back, and close the bug for good with a regression test that fails on the old code and passes on the new.55 minIntermediateHostedPro# debug-agent · agentPOST /api/agent/invoke200 OK · graded
Memory, protocols and many agents
Persist state, speak MCP from both sides, connect agents over A2A, and add planners, critics and human approval.
- Lab 12Add Long-Term Memory to an AI Agent: LangGraph + MilvusBuild a sales assistant that remembers: short-term state in a LangGraph checkpointer, long-term facts in Milvus, and reflection loops that extract knowledge.35 minIntermediateHostedPro# Milvus + LangGraph$ checkpoint.saveshort_term ...... 12 msgslong_term ....... 84 factsrecall@5 = 0.92
- Lab 13Build an MCP Tool Server & Connect a LangChain AgentBuild a Model Context Protocol server that exposes your company's tools and data, then connect a LangChain agent to it and see how MCP decouples tools from agents.40 minAdvancedHostedPro
- Lab 14Build an MCP Client and Host: Handshake, Tool Discovery, Error Handling and ApprovalsWrite the host side of the Model Context Protocol from scratch. Speak JSON-RPC over stdio to two real MCP servers, handle notifications, server requests, timeouts with cancellation and dead servers, discover and namespace their tools, turn every failure into text a model can use, run a tool loop with Llama 3.3 70B, and decide which calls need a person's approval.60 minIntermediateHostedPro# mcp-client · agentPOST /api/agent/invoke200 OK · graded
- Lab 15Build Two Agents That Talk via the A2A ProtocolBuild two independent agents that talk over the A2A protocol, each in its own process and found through an AgentCard, and see when A2A beats one orchestrator.40 minAdvancedHostedPro
- Lab 16Build a Multi-Agent Supervisor with LangGraphBuild a supervisor agent that routes queries to specialist agents, the core orchestration pattern the NCP-AAI exam tests.40 minIntermediateHostedPro
- Lab 17Planner-Executor Agents with Self-Critique: Plans, Validators, Replanning and Stopping RulesBuild a trip-planning agent that plans its searches up front, runs them in code, lets the model choose by id while code computes, checks every policy rule with a validator, replans when the solver says what is missing, stops for clear reasons, and measures whether self-critique earns its extra model calls.60 minIntermediateHostedPro# planner-executor-critique · agentPOST /api/agent/invoke200 OK · graded
- Lab 18Human-in-the-Loop Agent: Approval Interrupts, Durable Pauses, Crash-Safe Tools and EscalationBuild an accounts-payable agent that stops for a person before risky actions and survives the wait. Write the approval policy, checkpoint runs to disk and pause on approval requests, resume with approve, edit or reject decisions, journal tool calls so a crash can never pay twice, and expire approvals nobody decides.60 minIntermediateHostedPro# human-in-the-loop-agent · agentPOST /api/agent/invoke200 OK · graded
Evaluate and control it
Judge trajectories, optimise prompts with DSPy, route between models, and gate tool calls before they run.
- Lab 19Evaluate an Agent with LLM-as-JudgeBuild an eval harness that scores agent responses automatically: a reference-based judge for correctness, an accuracy metric and A/B comparison, the pattern NeMo Evaluator uses.30 minIntermediateHostedPro# agent-evaluation · agentPOST /api/agent/invoke200 OK · graded
- Lab 20Evaluate an AI Agent: Trajectories, Tool Calls and an LLM JudgeBuild the evaluation harness for a tool-using agent: record its trajectories on a golden set, grade final answers deterministically, score tool trajectories with exact, in-order and any-order matching plus precision and recall, check tool arguments, add an LLM judge with an order-swapped pairwise mode, measure the judge's agreement with human labels, run the suite into a report, and gate a prompt change on per-case regressions rather than the aggregate score.80 minIntermediateHostedPro# agent-evaluation-harness · agentPOST /api/agent/invoke200 OK · graded
Lab 21Model Routing & Cost Cascade with NIMSave 60 to 80% on inference by cascading queries from cheap to mid-size to large NIM models, and measure the real cost with NIM's usage.cost field against an always-large baseline.25 minIntermediateHostedPro
Lab 22Optimise Prompts with DSPy: Signatures, Metrics, Few-Shot Compilers and Held-Out TestsDeclare a ticket-triage task as a DSPy signature, score it with a metric, read the misses, compile it with labelled and bootstrapped demos, search demo sets on dev and catch the flattering score on a held-out test, then add the notes no label contains.60 minIntermediateHostedPro
Lab 23Gate an Agent's Tool Calls with JevBuild the guardrail that sits between a coding agent and its tools. Ask Jev, TypeSafe's decision model, typed questions about every proposed shell command, file write and HTTP request, turn its probabilities into allow, ask or block, measure the gate on labelled calls, and add the rules in code that catch what the model cannot see.45 minBeginnerHostedFree
Lab 24Route Between a Small and a Large Model with JevBuild the router that decides, per request, whether a cheap fast model can answer or the large model has to. Ask Jev, TypeSafe's decision model, whether a support request combines policy rules, replay recorded answers from GLM-5.3-flash and GLM-5.3 to measure accuracy and cost, tune the threshold on a dev split, confirm it on a test split, and make the router safe for production.60 minIntermediateHostedPro
Lab 25Calibrate Jev's Probabilities and ThresholdsTurn Jev's probabilities into numbers you can set a policy on. Build the reuse check of a semantic cache, measure how well Jev's probabilities match reality with reliability tables, ECE and Brier score, recalibrate them with Platt scaling, pick a threshold for a precision target on a dev split and confirm it on a test split, then test whether a different question wording really ranks better.60 minIntermediateHostedPro
Lab 26Find Where Jev Fails Before Your Users DoBuild the probes that tell you whether a question is safe to hand to Jev, TypeSafe's decision model. Catch facts missing from the state, confident answers to arithmetic that code does exactly, closed-world choice questions that force off-topic input into a real option, and instability under reordered options or text that argues with the model, then turn the results into a ship or fix verdict per feature.50 minIntermediateHostedPro
Attack and defend agents
Three labs from the AI security track hijack an agent through its tools and permissions, then re-scope it to least privilege.
Graded project
Build a tool-using ReAct agent
Build a ReAct agent that calls real tools on your own, then submit it for a score on every rubric criterion.
Guides for this track
Related collections:NVIDIA NIM vs NeMo
Questions about this track
Yes. The agent tool loop written without a framework and the Jev tool-call gate are free. The rest of the track is part of Pro.
No. The first lab writes the agent loop with no framework at all, and the LangGraph labs introduce nodes, edges and state as they appear.
Every step runs a check against your code and the running agent, and tells you what is still wrong before you move on.
Agent architecture, tool use, memory, multi-agent orchestration and evaluation are core to NVIDIA NCP-AAI and to the Anthropic Claude Certified Architect exams.
Other tracks
RAG and search
Chunking, hybrid search and reranking, permission-aware retrieval, GraphRAG and RAG evaluation.
LLMOps and MLOps
vLLM serving, load tests against SLOs, tracing, drift monitoring and prompt tests in CI.
AI security and red teaming
Prompt injection, tool poisoning and data exfiltration against live targets, then the defenses.
Every lab with Pro
This track and every other one, plus every practice test. $29.99 a month, cancel any time.