Working through realistic scenarios is the fastest way to find out whether you can pass the NVIDIA-Certified Professional: Agentic AI (NCP-AAI) exam or whether you only recognize the vocabulary. This set gives you 20 fresh practice questions written to the exam's own shape: a production situation with constraints, four options that are all real techniques, and one decision to make (five of the questions ask you to select two). Every domain of the 10-domain blueprint appears at least once, in rough proportion to its weight, and each question ends with an explanation that says why the winning option wins and why each distractor loses. Score yourself with the rubric below, then use the domain map at the end to decide what to read next.
Start Here
New to the exam? Read the NCP-AAI complete guide for the format, the ten domains, and a study path, and keep the NCP-AAI cheat sheet open while you review the explanations. When you are ready for full-length timed practice, the seven NCP-AAI practice tests are at /certificates/agentic-ai-professional, and you can try the free NCP-AAI sample questions first.
How NCP-AAI questions are built, and how to score yourself
NVIDIA publishes the shape of the exam: 60 to 70 questions in 120 minutes, delivered online with remote proctoring, ten domains weighted from 15 percent down to 5 percent, and no published passing score. That works out to a little under two minutes per question, which matters because NCP-AAI stems are scenarios rather than definitions. A typical stem names a system (a support agent, a RAG pipeline, a NIM deployment), states two or three constraints (a latency budget, a compliance rule, a cost ceiling, a preference for fewer moving parts), and asks for the best next step. The options are usually all legitimate techniques; the wrong ones solve a different problem, ignore a stated constraint, or cost more than the situation justifies. Multiple-response items ("Select TWO") appear regularly and require both correct choices.
The 20 questions below follow that construction. Answer all of them before reading any explanation, note the domain of every miss, and score yourself with this rubric.
Self-check rubric for this set
| Score on this set | Reading | What to do next |
|---|---|---|
| 17 to 20 correct | Strong | Move to timed full-length tests and work on pace; revisit only the domains you missed here |
| 14 to 16 correct | Borderline | Reread the explanation for each miss, then study the mapped articles for those domains before your next full test |
| Under 14 correct | Study first | Work through the complete guide and cheat sheet domain by domain, then return to this set before attempting timed tests |
Preparing for NCP-AAI? Practice with 455+ exam questions
Domain coverage in this set
The distribution below mirrors the published weights as closely as 20 questions allow. Five questions (2, 6, 11, 15, and 18) are multiple-response.
NCP-AAI domains and where they appear below
| Domain | Exam weight | Questions in this set |
|---|---|---|
| Agent Architecture & Design | 15% | 1, 2, 3 |
| Agent Development | 15% | 4, 5, 6 |
| Evaluation & Tuning | 13% | 7, 8, 9 |
| Deployment & Scaling | 13% | 10, 11, 12 |
| Cognition, Planning & Memory | 10% | 13, 14 |
| Knowledge Integration & Data Handling | 10% | 15, 16 |
| NVIDIA Platform Implementation | 7% | 17 |
| Run, Monitor & Maintain | 5% | 18 |
| Safety, Ethics & Compliance | 5% | 19 |
| Human-AI Interaction & Oversight | 5% | 20 |
The questions
Questions 1–5
0/5 answeredA retail company is building a support assistant that must look up order status in a database, search a returns-policy knowledge base, and, when a refund is warranted, call a payments API. The latency budget is four seconds per turn, and the team wants the fewest moving parts that still handle all three capabilities reliably. Which architecture should you recommend?
You are designing a market-research system in LangGraph (a framework that models an agent workflow as a graph of nodes and edges over shared state) where a planner decomposes a brief into sub-questions, several researcher agents work concurrently, and a writer assembles the report. Reviewers report that researchers duplicate each other's work and the writer sometimes starts before all research is complete. Select TWO changes that address these problems directly.
An insurer runs six independently deployed agents (intake, fraud screening, pricing, document extraction, notification, audit). Each agent calls the others over point-to-point HTTP, and adding a compliance agent last quarter required code changes in four existing services. The architecture team wants future agents to react to relevant events without modifying existing ones. Which coordination pattern fits this requirement?
A logistics agent built with LangChain calls a reschedule_shipment tool that takes a shipment ID and an ISO-8601 date. In production, about one call in twenty fails because the model passes values like "next Tuesday" or omits the shipment ID, and the downstream API answers with HTTP 400. You want to cut these failures without changing the API. What should you do first?
A field-service agent must accept a technician's photo of an equipment nameplate together with a typed fault description, extract the model and serial number, and check warranty status through an internal API. Nameplates vary in layout, font, and lighting. Which implementation handles this most reliably inside an agentic pipeline?
Architecture questions reward the simplest design that meets every stated constraint; the agent architecture design patterns guide walks through when a single tool-using agent is enough and when decomposition earns its coordination cost.
Questions 6–10
0/5 answeredYour LangGraph agent occasionally answers questions about competitor pricing (out of scope) and sometimes calls the send_email tool without the user asking. The model is a Nemotron instruct model served through NVIDIA NIM (containerized inference microservices with an OpenAI-compatible API). Before adding heavier machinery, you want fixes at the prompt and tool-definition layer. Select TWO changes that address both behaviors there.
A retrieval-augmented generation (RAG) assistant answers HR-policy questions from an internal document store. Users report confident answers that cite the wrong policy version. The team currently tracks only end-to-end answer accuracy on a 200-question golden set (a fixed, labeled evaluation dataset). Which additional metric would most directly locate the source of this failure?
You want to compare a new plan-and-execute prompt against the current ReAct prompt for a travel-booking agent. Product insists that no customer sees a degraded booking flow during the comparison, and finance wants a cost comparison measured on real production traffic. Which testing approach satisfies both constraints?
After upgrading the LLM behind a claims-triage agent, task completion on the golden set stayed flat, but the agent now selects the wrong tool on about 12 percent of multi-tool cases, up from 3 percent. Two sprints of prompt changes recovered only part of the gap. What is the most appropriate next optimization step?
A NIM-served LLM runs behind a Kubernetes deployment that autoscales on CPU utilization. During peaks, requests queue for 20 seconds while CPU stays near idle, so no new replicas start. GPU nodes need about 90 seconds to become ready. Which change most directly fixes the missed scale-out?
For Evaluation and Deployment questions, decide which layer failed (retrieval, generation, tool routing, or infrastructure) before you pick a fix; the agent evaluation and performance metrics guide maps each failure type to the metric that exposes it.
Tap each span in the trace (retrieval, prompt assembly, the model call, the tool call) and note what it captured: Questions 7 and 9 are answered by naming which of these stages failed before you pick a fix.
Questions 11–15
0/5 answeredYou are shipping a new version of a stateful customer-service agent whose conversation state is checkpointed to Postgres through a LangGraph checkpointer (a component that persists graph state after each step so a thread can resume). The new version changes the state schema. Operations wants zero dropped conversations and a fast rollback path if quality regresses. Select TWO practices that meet both requirements.
A RAG agent stack runs three models on one eight-GPU node: a 70B-class LLM, an embedding model, and a reranker. The LLM saturates its GPUs at peak while the embedding and reranking services sit mostly idle on their own dedicated GPUs. You must raise LLM throughput without buying hardware. Which change is the best first step?
An agent generates weekly staffing schedules subject to labor rules, employee preferences, and coverage minimums. Early attempts commit to a first assignment and later hit dead ends, producing invalid schedules. Latency is flexible (minutes are acceptable) and correctness matters more than token cost. Which reasoning pattern best fits this problem?
A concierge agent holds multi-week conversations with returning travelers. Two problems appear: long threads exceed the model's context window mid-trip, and preferences a traveler stated in March (aisle seat, vegetarian meals) are forgotten by June. The team wants both fixed with minimal added latency per turn. Which memory design should you adopt?
Engineers query a RAG assistant over 40,000 pages of equipment manuals. Dense semantic retrieval performs well on descriptive questions but misses exact part numbers such as "XR-2210-B", returning chunks about similar-looking parts instead. Top-k is 5 and there is no reranking stage. Select TWO changes that most directly improve retrieval for these exact-match queries.
Memory and RAG questions usually hinge on separating what belongs in the prompt from what belongs in a store; see memory management patterns for AI agents and the RAG systems and knowledge integration guide.
Quick check
You are ingesting quarterly financial PDFs into a RAG pipeline. Analysts complain that questions about figures in tables return prose from nearby pages, and charts are ignored entirely. The current pipeline extracts plain text and splits it into fixed 512-token chunks with no element detection. Which change addresses the root cause?
The 5 percent domains (monitoring, safety, human oversight) are short on weight and long on precise vocabulary; the safety guardrails guide and the observability and monitoring guide cover the rail types, span semantics, and approval patterns these questions test.
Master These Concepts with Practice
Our NCP-AAI practice bundle includes:
- 7 full practice exams (455+ questions)
- Detailed explanations for every answer
- Domain-by-domain performance tracking
30-day money-back guarantee
If you missed these, read this
Use your misses to pick reading. Each line names the questions for a domain and the sibling articles that teach the underlying pattern.
- Agent Architecture & Design (Questions 1 to 3): agent architecture design patterns and multi-agent collaboration essentials.
- Agent Development (Questions 4 to 6): tool use and function calling, prompt engineering best practices, and error handling and resilience patterns.
- Evaluation & Tuning (Questions 7 to 9): agent evaluation and performance metrics and testing strategies for agentic applications.
- Deployment & Scaling (Questions 10 to 12): NIM deployment strategies and Triton Inference Server for agentic workloads.
- Cognition, Planning & Memory (Questions 13 and 14): planning strategies: ReAct, Chain-of-Thought, Tree-of-Thoughts and memory management patterns.
- Knowledge Integration & Data Handling (Questions 15 and 16): RAG systems and knowledge integration and vector databases for agentic AI.
- NVIDIA Platform Implementation (Question 17): integrating NVIDIA NIM with LangChain.
- Run, Monitor & Maintain (Question 18): agent observability and monitoring.
- Safety, Ethics & Compliance (Question 19): safety guardrails for agentic AI and ethics and compliance.
- Human-AI Interaction & Oversight (Question 20): the HITL escalation section of the NCP-AAI cheat sheet.
When the misses cluster in Evaluation & Tuning, Agent Architecture & Design, or Cognition, Planning & Memory, you are in good company: those are the domains where Preporato users score lowest, and the hardest NCP-AAI topics analysis breaks down the specific question patterns behind that.
Evaluate an agent with an LLM judge, then add memory
Build an LLM-as-judge harness that scores agent runs and a LangGraph agent that keeps long-term facts in Milvus, the two systems behind the Evaluation and the Cognition, Planning and Memory questions most readers miss.
Frequently asked questions
Key takeaways
Key Takeaways
0/8 completedNext steps
Take the 20 questions again in a week without looking at the answers; a second pass shows whether you learned the pattern or memorized the letter. Then move to full-length timed practice on the NCP-AAI certificate page, starting with the free sample questions if you want to see the interface first. All seven NCP-AAI tests, plus the hands-on labs that mirror the deployment and RAG scenarios above, are included in Preporato Pro; details are on the pricing page.
Build the RAG pipeline and ReAct agent behind these stems
Stand up retrieval over NVIDIA NIM endpoints and run a ReAct loop against them, so the deployment, retrieval, and platform scenarios in this set become systems you have already operated.
Sources:
- NVIDIA-Certified Professional: Agentic AI (NCP-AAI) certification page
- NVIDIA NIM documentation
- NVIDIA NeMo Guardrails documentation
- NVIDIA NeMo Retriever
- LangGraph documentation
Ready to Pass the NCP-AAI Exam?
Join thousands who passed with Preporato practice tests
![NCP-AAI Practice Questions With Explanations: 20 Scenarios [2026]](/blog/ncp-aai-practice-questions-with-explanations-2026.webp)