AI Engineer course, built on hands-on GPU labs
From transformers to production-grade RAG agents. Hands-on, GPU-backed.
What you'll be able to build
Eight capabilities, previewed with the actual animations from each module.
Build a transformer from scratch. Attention, multi-head, residuals, and the full block, trained on real data.
Open this moduleAbout this course
An AI Engineer is the role that ships LLM-powered features to production: choosing the right model, fine-tuning when off-the-shelf isn't enough, building retrieval and tool-use loops, and running it all on GPUs that don't time out. This path teaches the role through real labs. Every concept ends in code that runs against a real GPU.
Skills you'll put on a resume
- Implement a decoder-only transformer end-to-end and train it on real data
- Fine-tune open-weight LLMs with LoRA and QLoRA on a single GPU
- Build a production RAG pipeline with hybrid search, reranking, and grounded generation
- Serve LLMs at production throughput with vLLM (PagedAttention, continuous batching, prefix caching)
- Evaluate models with perplexity, LLM-as-judge, and preference data
- Build production agents with ReAct, MCP tool servers, multi-agent orchestration, and memory
- Pass the NVIDIA NCA-GENL, NCP-GENL, and NCP-AAI certifications
For
Software engineers comfortable with Python and basic ML who want to move into LLM/AI engineering work
Prerequisites
- Comfortable Python (functions, classes, package management)
- Familiarity with PyTorch or another ML framework
- Basic ML literacy (training/eval, gradient descent, overfitting)
Every lab in this course, module by module
What you break, then what you fix
01How LLMs actually work
For software engineers with zero AI background. You'll come out the other side able to hold a conversation about transformers, read a model card without bouncing, and reason about why a model just did what it did. Every term gets defined inline; we lead from software-engineering mental models (APIs, tokens-as-bytes, scheduling) before touching any ML jargon. The labs at the end of the module put the theory through real weights.
- Data Preparation for LLM Training45 min · intermediateBuild a real pretraining/instruction data pipeline: load a raw corpus, apply quality filters, deduplicate, train a BPE tokenizer, and batch-validate on GPU. This is the unglamorous work that actually decides how good your model will be.
- Build a Transformer from Scratch: Attention, Masking & LayerNorm50 min · advancedBuild every piece of a decoder-only transformer by hand — scaled dot-product attention, multi-head attention, the full block with residuals and LayerNorm, then assemble a tiny GPT and train it. No shortcuts, no pre-built attention modules.
- Train a Small Language Model from Scratch55 min · advancedTrain a real GPT-style language model from zero on TinyStories: tokenize, wire up the optimizer and LR schedule, run the training loop with validation perplexity, and generate coherent text from your own weights. End-to-end pretraining in minutes on one GPU.
02Using LLMs via API
Build your first working LLM feature against a hosted API. You'll structure prompts, count tokens before you call, control cost with caching, harden against prompt injection, and know when prompting alone stops being enough. Every concept lands in code you'd actually write against OpenAI, Anthropic, or an open-weight model behind vLLM.
Read first: How to Build an AI Agent From Scratch: The Tool Loop, No Framework, LLM Gateway Guide: What It Does, What to Build Yourself, and When to Buy One
- Build an Agent From Scratch: The Tool Loop Without a Framework80 min · intermediateHand-roll a tool-using agent on the chat completions API and nothing else: parse the model's tool calls, dispatch them to plain Python functions, close the loop with correctly threaded tool results, add stop conditions and budgets, turn every failure into a tool result the model can read, put an approval guardrail in front of a side-effecting refund tool, control context with truncation and a trace, and finish with an evaluation run that scores the agent on a fixed question set.
- Build an LLM Gateway: One Door Between Your Apps and the Model90 min · intermediatePut a gateway you wrote in front of the model and give it the four floor capabilities every LLM feature needs: API keys with a request log, a price on every call, per-app rate limits and daily budgets, an exact-match response cache, retries with backoff and idempotency keys, model routing with a fallback, and a report over a replayed morning of traffic that shows who spent what.
- Ship a Production LLM API Featuregraded project
- Build a Structured-Output Extraction Servicegraded project
03LLM Inference & Optimization
Loading a model is easy; serving it without going broke isn't. Learn the precision/throughput/VRAM tradeoffs by quantizing models, sweeping batch sizes, and building a mini-Triton with dynamic batching.
- Quantize & Optimize LLMs with bitsandbytes40 min · intermediateLoad a model in fp16, INT8, and NF4, then benchmark the three precisions on VRAM, latency, and output quality. See where quantization wins and where it costs you.
- Batch Size & Precision Sweep: Finding Your Sweet Spot40 min · intermediateSweep batch sizes and numerical precisions (fp32, fp16, bf16) on a real model to find the throughput/VRAM knee, then ship a production recommendation with SKU-aware precision picks and an accuracy gate.
- Inference Serving Patterns: Dynamic Batching, Throughput, and the Triton Mental Model40 min · intermediateBuild a mini-Triton inference server in ~30 lines of Python: a dynamic batcher with max_batch_size and max_queue_delay knobs, load-tested against a naive baseline, swept for the throughput-latency tradeoff, and bridged to a real Triton config.pbtxt.
04Retrieval-Augmented Generation
RAG is the single most-shipped LLM pattern in production. Build it end-to-end: embeddings, vector search, hybrid retrieval (dense plus BM25), reranking with a cross-encoder, and grounded generation. Then build the same pipeline on NVIDIA NIM.
Read first: LLM Evaluation Metrics: Which Number to Use for Which Failure
- Retrieval-Augmented Generation (RAG) Pipeline with Local Models45 min · intermediateBuild an end-to-end RAG pipeline on a single GPU: BGE embeddings, L2-normalized vector retrieval by dot product, and a local generator that answers with and without retrieved context so you can see exactly what retrieval changes.
- Advanced RAG: Hybrid Search + Cross-Encoder Reranking40 min · advancedBuild a production-shape retrieval stack — dense bi-encoder plus from-scratch BM25, fused with Reciprocal Rank Fusion, then re-ordered by a BAAI cross-encoder. The exact architecture behind modern enterprise RAG.
- Build a RAG Pipeline with NVIDIA NIM35 min · intermediateBuild a complete Retrieval Augmented Generation pipeline — from document chunking to vector search to an agent that answers questions from your knowledge base.
- Evaluate a RAG Pipeline: Retrieval Metrics, Faithfulness and a Regression Gate80 min · intermediateTurn 'the RAG bot seems better' into numbers you can gate on: chunk a knowledge base with paragraph provenance, label a golden set, score retrieval with recall, precision, MRR and nDCG, sweep chunk sizes, judge answer faithfulness claim by claim with a model, grade correctness, triage every failure by where it happened, and block a candidate configuration that regresses a protected question.
- Build a LangChain or LlamaIndex RAG Pipelinegraded project
- Build a RAG-Powered Support Assistantgraded project
05Fine-Tuning & Alignment
Fine-tuning is how generic models become useful for your domain. LoRA on Llama 3, continued pretraining on a code corpus, and DPO alignment with preference data: three different ways to move a model's behavior, each with its own cost and risk profile.
- Fine-Tune an LLM with LoRA and QLoRA (Jupyter)45 min · intermediateFine-tune Meta Llama 3 8B on a custom instruction dataset using LoRA and QLoRA. Learn parameter-efficient fine-tuning from data preparation through evaluation — the #1 most demanded AI skill.
- Continued Pre-Training: Adapt a Pretrained LM to a New Domain45 min · advancedTake GPT-2 and domain-adapt it to Python code in 150 steps, measuring both the gain on code and the cost in catastrophic forgetting on English. The exact recipe behind Code Llama, BloombergGPT, and every domain-specialized LLM of the last three years.
- RLHF & DPO Alignment55 min · advancedRun real Direct Preference Optimization on a small language model with TRL's DPOTrainer. Capture a baseline, build a preference dataset, train, and measurably shift the model's behavior in four steps.
- Synthetic Data Generation for Model Training40 min · intermediateBuild a Self-Instruct style synthetic dataset end-to-end: seed instructions, LLM-driven generation, robust parsing, quality filtering, and dedup + diversity scoring. The same pipeline that produced Alpaca, WizardLM, and most modern instruction-tuning corpora.
06Production Serving
Take a fine-tuned model from notebook to production. Stand up vLLM with PagedAttention and continuous batching, package the runtime as a versioned GPU container, and ship it through a CI pipeline with rollback safety.
- vLLM Production Serving: PagedAttention, Continuous Batching, Prefix Caching55 min · advancedStand up vLLM and measure the three features that make it the de-facto inference server: PagedAttention's KV-cache capacity, continuous batching throughput, and prefix caching speedups. Then write the production spec — server args, Kubernetes deployment, monitoring, autoscaling.
- Deploy & Serve LLMs in Production (Jupyter)45 min · intermediateGo from slow single-request inference to production-ready LLM serving with vLLM. Benchmark throughput, tune settings, and learn when to use vLLM vs Triton vs TGI.
- GPU Container Lifecycle: Build, Test, Ship, Rollback40 min · intermediateWalk through the full lifecycle of a production GPU container — multi-stage Dockerfile, self-hosted GPU CI, a fail-fast smoke test, and a Kubernetes Deployment with readiness probes gated on real GPU compute. The pipeline that stops bad images before users see a 500.
07Evaluation & MLOps
An LLM you can't measure is one you can't trust. Run perplexity, BLEU, and LLM-as-judge with position-bias detection, then wire experiment tracking with MLflow so every run is reproducible and registered.
Read first: LLM Observability: What to Trace, What to Measure and What to Keep, AI Agent Evaluation: Trajectories, Tool Calls and an LLM-as-a-Judge, LLM-as-a-Judge: How to Use a Model to Grade Model Outputs Without Fooling Yourself, LLM Evaluation Metrics: Which Number to Use for Which Failure
- Evaluate an AI Agent: Trajectories, Tool Calls and an LLM Judge80 min · intermediateBuild the evaluation harness for a tool-using agent: record its trajectories on a golden set, grade final answers deterministically, score tool trajectories with exact, in-order and any-order matching plus precision and recall, check tool arguments, add an LLM judge with an order-swapped pairwise mode, measure the judge's agreement with human labels, run the suite into a report, and gate a prompt change on per-case regressions rather than the aggregate score.
- LLM Observability: Trace, Cost and Debug a RAG Assistant75 min · intermediateInstrument a real RAG assistant with traces from scratch: nested spans around retrieval and generation, token usage and cost attributed to every call, a latency and cost report, a triage tool that finds retrieval misses from trace evidence, PII redaction before anything is written, sampling that keeps every error and slow request, a cost budget alarm, and a CI gate that blocks a prompt change that makes the assistant slower or dearer.
- Evaluation & Benchmarking LLMs45 min · intermediateFour evaluation lenses in one lab: compute real perplexity, expose BLEU's blindness to paraphrase, run side-by-side model comparisons, and build an LLM-as-judge harness with position-bias detection.
- MLflow Experiment Tracking: From Single Run to Team Workflow35 min · intermediateWire the four load-bearing pieces of MLflow into a real training loop — tracked runs with params and metrics, a registered model with stage transitions, a multi-run sweep + search, and a production spec (server, k8s Job, tags, autolog).
- Build an LLM Evaluation Harnessgraded project
08Agents & Tool Use
An LLM that can't call tools is a debate partner. An LLM that can call tools is an employee. Build a ReAct agent on NVIDIA NIM, compare three orchestration patterns (ReAct vs tool-calling vs plan-and-execute), expose your tools via MCP, route queries with a supervisor across specialist agents, add short and long-term memory, harden it with guardrails, and put the whole thing under an LLM-as-judge eval harness. The full agentic stack the NCP-AAI exam tests and that production teams actually ship.
Read first: Agentic AI Course Guide: What a Good One Makes You Build, How to Build an AI Agent From Scratch: The Tool Loop, No Framework
- Build a ReAct Agent with NVIDIA NIM35 min · intermediateBuild an AI research librarian: an agent that searches a corpus of ML papers, compares methods and answers multi-step questions on real NIM endpoints with LangGraph.
- Build an AI Agent 3 Ways: ReAct vs Tool Calling vs Plan-and-Execute35 min · intermediateBuild the same SaaS support agent three ways, as ReAct, direct tool calling and plan-and-execute, then compare speed, reasoning quality and reliability to learn when each fits.
- Build an MCP Tool Server & Connect a LangChain Agent40 min · advancedBuild a Model Context Protocol server that exposes your company's tools and data, then connect a LangChain agent to it and see how MCP decouples tools from agents.
- Build a Multi-Agent Supervisor with LangGraph40 min · intermediateBuild a supervisor agent that routes queries to specialist agents, the core orchestration pattern the NCP-AAI exam tests.
- Add Long-Term Memory to an AI Agent: LangGraph + Milvus35 min · intermediateBuild a sales assistant that remembers: short-term state in a LangGraph checkpointer, long-term facts in Milvus, and reflection loops that extract knowledge.
- Build NeMo Guardrails for an AI Agent: Jailbreak & Topical Rails35 min · intermediateBuild a guarded IT support agent that blocks jailbreaks, refuses off-topic questions and handles IT queries safely, using keyword checks, LLM validation and NeMo Guardrails.
- Evaluate an Agent with LLM-as-Judge30 min · intermediateBuild an eval harness that scores agent responses automatically: a reference-based judge for correctness, an accuracy metric and A/B comparison, the pattern NeMo Evaluator uses.
- Build a Tool-Using ReAct Agentgraded project
Lab paths through this course
The same labs, grouped by topic
- RAG pipeline course: build, evaluate and secure RAGSeven worked examples, from your first vector search to a RAG firewall.
- Agentic AI course: build AI agents hands-onEleven worked examples, from a tool loop written by hand to a guarded, evaluated multi-agent system.
- How to fine-tune an LLM: LoRA, QLoRA and DPO on real GPUsEight GPU labs from data preparation to a model you trained yourself.
- LLM inference optimization course: vLLM, quantization and profilingNine labs on the numbers that decide whether an LLM deployment is affordable: VRAM, latency, throughput, dollars.
- NVIDIA NIM vs NeMo: the difference, then build with bothEight hosted labs on real NIM endpoints, from a ReAct agent and a free RAG build to NeMo Guardrails.
- How to train an LLM from scratch: tokenizer to pretraining on real GPUsEight GPU labs from a raw corpus to coherent text generated from weights you trained yourself.
Guides & articles
Deep-dive reading that pairs with this course
AI Engineer Roadmap 2026: From Python to Production LLM Systems
A stage-by-stage AI engineer roadmap for 2026: what to learn and build at each stage, where NVIDIA certifications fit, and how the path plays out in India.
ReadHow to Become an AI Engineer: A Practical Guide for Working Developers
How to become an AI engineer from a software, data, or student background: the skills that get you hired, a 90-day build plan, and how interviews test you.
ReadAI Engineer Certifications in 2026: Which Ones Are Worth the Fee
AI engineer certifications in 2026, verified on vendor pages: AWS, Google, Microsoft, Databricks, and NVIDIA exams, with costs, formats, and which fits you.
ReadAI Engineer Skills and Salary in 2026: What Employers Pay For
AI engineer salary in 2026 from Built In and Glassdoor for the US and India, why the sources disagree, and the eight skills that move the number.
ReadAI Engineer vs ML Engineer: Roles, Skills, and Which to Pursue
AI engineer vs ML engineer by what each ships: predictive models on your data versus products on foundation models. Skills side by side, overlap, how to choose.
ReadMLOps Course Guide: What a Good One Teaches and What You Should Ship by the End
What an MLOps course must teach: the training-to-monitoring loop, the eight units that cover it, five things to ship by the end, and a six-week syllabus.
ReadMLOps Certifications in 2026: Options, Costs, and What They Actually Test
MLOps certifications in 2026, verified on vendor pages: Databricks, AWS, Google Cloud, and NVIDIA exams with costs, formats, MLOps weight, and the retired one.
ReadMLOps vs DevOps (and AIOps): What Changes When the Artifact Is a Model
MLOps vs DevOps stage by stage: what stays the same, what changes when the artifact is a model, where AIOps fits, and what each engineer must add to cross over.
ReadLLM Observability: What to Trace, What to Measure and What to Keep
A hands-on guide to LLM observability: the trace schema, the signals worth a dashboard, cost per answer from the usage block, sampling that keeps every failure, redaction, and a regression gate for prompt changes.
ReadHow to Build an AI Agent From Scratch: The Tool Loop, No Framework
Build an AI agent on the chat completions API alone: parse tool calls, dispatch to plain functions, thread results back, add budgets, handle malformed arguments, guard side effects, and evaluate the result.
ReadAI Agent Evaluation: Trajectories, Tool Calls and an LLM-as-a-Judge
How to evaluate an AI agent: a golden set with expected answers, tool paths and arguments, deterministic graders, an LLM-as-a-judge checked for position bias and calibrated against human labels, and a per-case regression gate.
ReadLLM-as-a-Judge: How to Use a Model to Grade Model Outputs Without Fooling Yourself
A practical guide to LLM-as-a-judge: when a judge belongs in an evaluation, how to write the rubric and the output format, the biases that skew verdicts and the order swap that exposes them, and how to calibrate the judge against human labels before trusting it.
ReadAgentic AI Course Guide: What a Good One Makes You Build
What an agentic AI course must cover and what you should have shipped by the end: the tool loop, budgets and error handling, tool design, memory, agent patterns, MCP, multi-agent routing, guardrails, evaluation and observability, with a self-study syllabus.
ReadLLM Gateway Guide: What It Does, What to Build Yourself, and When to Buy One
What an LLM gateway does in front of your model provider: keys and logging, pricing per call, rate limits and budgets, exact-match caching, retries and idempotency, routing with fallback, and the observability that makes the bill explainable, with a build-or-buy answer.
ReadLLM Evaluation Metrics: Which Number to Use for Which Failure
A working map of LLM evaluation metrics by what they grade: retrieval (recall, precision, MRR, nDCG), answers against a reference, answers against their context (faithfulness), agents by trajectory, judges by agreement, and operations by latency and cost, with the failure each one catches and the ways each one lies.
ReadReady to start?
Pro gives you all 31 labs in this path, every other lab on Preporato, and every practice test. $29.99/mo, cancel anytime.