Microsoft's study guide for AI-103 lists the exam's scope as short objectives such as "Configure model and agent deployments" or "Implement orchestrated multi-agent solutions". Each line stands for a set of specific product decisions, and the questions test those decisions. This page takes the skills measured as of April 16, 2026, area by area, and spells out what each objective means in Microsoft Foundry, which services and settings answer it, and where questions tend to set traps.
Find out where you stand first
Take the free AI-103 practice questions before you read on, then spend the most time on the areas where you dropped marks. The AI-103 course on preporato.com follows these five areas in the same order, one module per area.
How to read the skills list
Microsoft adds two notes that change how you should study. The bullets under each skill illustrate how it is assessed, and related topics can appear even when no bullet names them. Most questions cover generally available features, and preview features can appear when they are commonly used. The English exam is updated first, and localized versions follow about eight weeks later, so check the study guide page for the date that applies to you.
Preparing for AI-103? Practice with 390+ exam questions
1. Plan and manage an Azure AI solution
Share of the exam: 25 to 30 percent. This area covers the decisions that come before any code and the controls that wrap around everything afterwards.
Choose the appropriate Foundry services for generative AI and agents
Choosing services
| Objective | What it means in practice |
|---|---|
| Choose a model for each task | Large models for open reasoning, small models for cost and latency, multimodal models for images and audio, and Foundry Tools when a fixed-output task such as PII detection or translation has a purpose-built service |
| Choose services for generation, grounding, vector search, agent workflows or multimodal processing | Foundry Models for generation, Azure AI Search and Foundry IQ for grounding and vectors, Foundry Agent Service and Microsoft Agent Framework for agents, Content Understanding for documents and media |
| Choose a retrieval and indexing method | Keyword, vector, hybrid or semantic ranking, integrated vectorization during indexing, and agentic retrieval when an agent needs a knowledge base |
| Choose memory, tool and knowledge services for agents | Conversation state in the Responses API, the memory tool, function and OpenAPI tools, MCP servers, toolboxes, and knowledge bases |
Set up AI solutions in Foundry
Setting up
| Objective | What it means in practice |
|---|---|
| Design infrastructure for AI apps and agents | A Foundry resource that hosts projects, with network isolation, a managed identity, and the standard agent setup when agent data must stay in your own Cosmos DB, Storage and Azure AI Search resources |
| Choose deployment options | Global, Data Zone or Regional deployments, standard per-token or provisioned throughput, Batch for asynchronous work, and priority processing for latency-sensitive traffic |
| Configure model and agent deployments | Model versions and upgrade policies, tokens-per-minute quota, and agent versions behind the endpoint that consumers call |
| Integrate Foundry projects with CI/CD | Pipelines that create projects, deploy models and agents, and grant each project identity its runtime roles |
Manage, monitor and secure AI systems
Running it safely
| Objective | What it means in practice |
|---|---|
| Manage quotas, scaling, rate limits and cost | Tokens-per-minute quota per deployment, 429 handling, spillover from provisioned to standard deployments, and reservations that match the deployment type |
| Monitor performance, drift, safety events and grounding quality | Azure Monitor metrics, traces, continuous and scheduled evaluations, and alerts |
| Monitor ingestion quality, index health and relevance | Indexer execution history, skill errors and warnings, and relevance checks on the search index |
| Configure security | Managed identities, private endpoints with linked private DNS zones, keyless access with Microsoft Entra ID, and least-privilege Foundry roles |
Implement responsible AI across generative AI and agentic systems
Responsible AI controls
| Objective | What it means in practice |
|---|---|
| Configure safety filters, guardrails, risk detection and moderation | Deployment guardrails with controls set to annotate or annotate and block at a severity threshold, Prompt Shields for direct and document attacks, and blocklists |
| Apply evaluators, safety evaluations and explanation tooling | Quality and safety evaluators, a judge deployment for AI-assisted ones, and the AI Red Teaming Agent for attack success rates |
| Implement auditing | Trace logging, provenance metadata such as Content Credentials, and approval workflows with recorded decisions |
| Govern agent behavior | Tool-access controls such as allowed_tools and require_approval on MCP tools, guardrail controls such as Task Adherence, and approval steps that keep a person in the loop |
Role questions are where small wording matters most. Foundry User is the least-privilege role for developers who build and test agents, and Foundry Agent Consumer is the one for principals that only call agents. An API key grants full access without role restrictions, so keyless access ends with key authentication disabled on the resource. Cost scenarios can describe spillover without naming it, as requests above provisioned capacity that move automatically to a standard deployment in the same resource. In safety scenarios, a control the platform enforces beats an instruction that asks the model to behave.
2. Implement generative AI and agentic solutions
Share of the exam: 30 to 35 percent. This is the largest area, and the platform has moved further here than anywhere else since older study material was written.
Build generative applications by using Foundry
Generative apps
| Objective | What it means in practice |
|---|---|
| Deploy and consume LLMs, small, code and multimodal models | Model deployments called through the Responses API on the stable v1 routes |
| Implement RAG in an application | Retrieve from Azure AI Search, pass the chunks as context, and cite the sources |
| Design workflows, tool-augmented flows and multistep reasoning | Function calling, reasoning models and Agent Framework workflows |
| Evaluate models and apps | Groundedness for fabrication, Relevance, Response Completeness against ground truth, and safety evaluators |
| Use Foundry SDKs and connectors | The azure-ai-projects client and the OpenAI client it returns for the same project endpoint |
| Connect an application to a Foundry project | The project endpoint, DefaultAzureCredential and a role on the project |
Build agents by using Foundry
Agents
| Objective | What it means in practice |
|---|---|
| Define roles, goals, conversation tracking and tool schemas | Instructions for a prompt agent, conversations or response chaining for state, and JSON schemas for function tools |
| Integrate retrieval, function calling and memory | Search or knowledge base tools, function tools that your code runs, and the memory tool |
| Integrate tools | OpenAPI tools, MCP servers, search, Content Understanding and custom functions, which a toolbox can expose to many agents behind one MCP endpoint |
| Implement multi-agent solutions | The sequential, concurrent, handoff, group chat and Magentic orchestrations in Microsoft Agent Framework |
| Build autonomous or semiautonomous workflows with approvals | require_approval on tools and human-in-the-loop steps that pause a workflow until a person decides |
| Monitor, evaluate and analyze errors | Traces, plus agent evaluators such as Intent Resolution, Task Adherence and Tool Call Accuracy |
Optimize and operationalize generative AI systems
Operating it
| Objective | What it means in practice |
|---|---|
| Tune generation behavior | Prompt engineering, sampling parameters, reasoning effort and structured outputs |
| Implement reflection, chain-of-thought evaluation and self-critique | A second pass that critiques and revises an answer, scored by evaluators |
| Set up observability | Traces that follow the OpenTelemetry GenAI conventions, token usage, safety signals and latency per step |
| Orchestrate multiple models or hybrid LLM and rules engines | Model router, which picks a model per request, and deterministic rules for steps that must not vary |
What does the agent remember between turns? Expect that question in several disguises. Chaining with previous_response_id only works when responses are stored, so with store set to false the history travels in the input array. Orchestration options each hang on a phrase: agents passing control to each other is handoff, and a manager that plans and replans is Magentic. For evaluation, check the inputs each evaluator needs, since Similarity and Response Completeness compare against ground truth while Groundedness checks against context.
3. Implement computer vision solutions
Share of the exam: 10 to 15 percent. Generation and editing sit next to analysis here, and the area closes with safety rules written for images and video.
Design and implement image and video generation solutions
Generation and editing
| Objective | What it means in practice |
|---|---|
| Generate images from text and reference media | GPT-image models, from a text prompt alone or with input images as references |
| Generate videos from text and reference media | Sora 2 in Foundry, with an optional reference image that anchors the first frame |
| Configure image editing | Inpainting through the edit API, with a same-size PNG mask whose transparent pixels mark the area to change, and input_fidelity to keep faces and style |
| Edit generated videos | Remix in Sora 2, which makes targeted changes to a finished video by its ID |
| Select generation and editing controls | Size, quality, count and output format for images, and size, length and reference inputs for video |
Design and implement multimodal understanding workflows
Understanding images and video
| Objective | What it means in practice |
|---|---|
| Analyze visual context with multimodal models | Images passed to a vision-capable model with the question |
| Caption single or multiple images, concise or detailed | Prompted captions with a length and detail target |
| Answer questions grounded in visual evidence | Answers that point at what the image shows |
| Generate alt text and extended descriptions | Text that follows accessibility guidance, short alt text with a longer description where needed |
| Use Content Understanding for visual characteristics and video | Image and video analyzers, segments described in natural language, and RAG-ready output from analyzers such as prebuilt-videoSearch |
| Identify objects, components or regions | Locating items in images and video frames |
One objective here still names "single-task and pro-mode Content Understanding pipelines". Pro mode only existed in the 2025-05-01-preview API, which has retired, and Microsoft points pro mode users to agentic mode in the 2026-06-01-preview API, a preview for document analysis.
Implement responsible AI for multimodal content
Visual safety
| Objective | What it means in practice |
|---|---|
| Filter unsafe or disallowed visual content | Azure AI Content Safety image analysis with severity levels, and custom categories for new symbols |
| Detect indirect prompt injection in images | Extract the text from the image and scan it with Prompt Shields as a document |
| Enforce visual policy rules | Watermarks and Content Credentials on generated images, and checks for prohibited symbols and brand misuse |
If an edit must leave most of an image untouched, the answer involves a mask, because a prompt alone leaves the whole image open to change. Text printed inside an image counts as third-party content, so the defense pairs detection on the extracted text with approval before any action that could do damage.
Master These Concepts with Practice
Our AI-103 practice bundle includes:
- 6 full practice exams (390+ questions)
- Detailed explanations for every answer
- Domain-by-domain performance tracking
30-day money-back guarantee
4. Implement text analysis solutions
Share of the exam: 10 to 15 percent. Language and speech both live here, with language models now doing much of the work next to the Foundry Tools.
Apply language model text analysis
Text analysis
| Objective | What it means in practice |
|---|---|
| Extract entities, topics, summaries and structured JSON | Prompting with structured outputs, or Azure Language in Foundry Tools for named entities and PII |
| Detect sentiment, tone, safety issues and sensitive content | Language models for tone, Content Safety for harm, and PII detection with a redaction policy for personal data |
| Translate text | Azure Translator in Foundry Tools, where the 2026-06-06 text API picks standard NMT or an LLM deployment with tone and gender controls, or a language model flow you build |
| Customize outputs for domain tasks | Prompts, examples and terminology for jobs such as compliance summaries |
Implement speech solutions
Speech
| Objective | What it means in practice |
|---|---|
| Convert speech to text and text to speech for agents | Azure Speech in Foundry Tools, with SSML to control pronunciation and pacing |
| Integrate speech as an agent modality | Voice Live, a speech-to-speech API over WebSockets for voice agents, plus phrase lists or custom speech for vocabulary |
| Enable reasoning from audio inputs | Audio-capable models that take speech directly |
| Translate speech | Speech translation in Azure Speech, or a language model step for the text |
Dates matter in this area. Several legacy Azure Language features, including sentiment analysis, key phrase extraction, summarization, conversational language understanding and custom question answering, retire on March 31, 2029, and Microsoft directs new projects to Foundry models. On the speech side, a short list of new names calls for a phrase list, which applies at runtime with no training, and a vocabulary of more than 2,000 phrases calls for custom speech.
5. Implement information extraction solutions
Share of the exam: 10 to 15 percent. It covers retrieval for grounding and extraction from documents, the two jobs that feed agents and RAG.
Build retrieval and grounding pipelines
Retrieval and grounding
| Objective | What it means in practice |
|---|---|
| Ingest and index documents, images, audio and video | Data sources, indexers on a schedule, skillsets and index projections that write one document per chunk |
| Configure semantic, hybrid and vector search | Hybrid queries merged with Reciprocal Rank Fusion, the semantic ranker reranking the top results, and vector fields with compression |
| Enrich with built-in or custom skills | OCR, Text Merge and Text Split skills, embedding skills, and custom Web API skills for your own code |
| Configure RAG ingestion with OCR | Normalized images extracted from documents, OCR, then merged text so image content becomes searchable |
| Connect retrieval to workflows and agent tools | The Azure AI Search tool for one index, or agentic retrieval through a knowledge base that agents call over MCP |
Extract content from documents
Document extraction
| Objective | What it means in practice |
|---|---|
| Combine OCR, layout analysis and field extraction | Content Understanding analyzers built on prebuilt-document, or domain analyzers such as prebuilt-invoice |
| Produce clean, grounded representations for agents and RAG | Markdown output with layout preserved, from analyzers such as prebuilt-documentSearch |
| Implement analyzers for structured or Markdown output | Custom analyzers with field schemas that extract, classify or generate values, plus confidence and source grounding |
Three settings catch people here. The vectorizer used at query time must match the embedding model used at indexing, dimensions included. Index projections left on the default mode add a parent document for each source file next to its chunks. And in Content Understanding, splitting a file that holds several document types takes content categories with enableSegment set to true, plus an analyzerId on any category whose segments need field extraction.
Turning the list into a study plan
Weight your time by the size of each area, and within an area, by how unfamiliar it is. Engineers with an AI-102 background usually need the most time on agents, guardrails and Content Understanding, and less on search or speech. For a hands-on build in each area and a plan for the exam day, read how to pass AI-103 on your first attempt.
Work through the AI-103 course, which has one module per area in this order, and test each area as you finish it. The six timed AI-103 practice exams report a score per skill area, so the next study block always targets the weakest one. For the wider picture, the AI-103 complete guide covers who the exam suits, and AI-103 vs AI-102 covers what changed since the previous exam.
Frequently asked questions
Sources:
- Study guide for Exam AI-103: Developing AI Apps and Agents on Azure
- What is Microsoft Foundry?
- Role-based access control for Microsoft Foundry
- Provisioned throughput for Foundry Models
- Azure OpenAI Responses API
- Microsoft Agent Framework orchestrations
- Image generation and editing in Microsoft Foundry
- Prompt Shields
- What is Azure Language?
- Improve recognition with phrase lists
- Define index projections in Azure AI Search
- Agentic mode in Content Understanding
- Content Understanding classifiers
Ready to Pass the AI-103 Exam?
Join thousands who passed with Preporato practice tests
