AI Infrastructure Engineer course, on real Kubernetes and real GPUs
Run GPU clusters that don't melt down. Kubernetes, GPU Operator, MIG, MPS, observability.
About this course
AI Infrastructure Engineers run the GPU clusters that everyone else's models depend on. The role lives at the intersection of Kubernetes platform engineering and GPU-specific operations, and almost no online course teaches it through real clusters. This path does. Every module is a lab on a live, isolated Kubernetes environment with real GPU scheduling, real GPU operator components, and real triage scenarios that break the way they break in production.
Skills you'll put on a resume
- Schedule GPU workloads on Kubernetes with the right resource requests, limits, and priority classes
- Install and triage the NVIDIA GPU Operator end-to-end, attributing every component's role
- Share single GPUs across workloads with CUDA streams, MPS, and MIG, and know when each fits
- Implement the four-stage cost-audit pipeline (measure → classify → price → recommend) for GPU fleets
- Profile PyTorch training with Nsight Systems and the built-in profiler to find real bottlenecks
- Size inference capacity on a real GPU with dynamic batching, vLLM PagedAttention and continuous batching, and a batch-size-by-precision sweep
- Pass the NVIDIA-Certified Associate AI Infrastructure & Operations exam
For
Platform engineers, SREs, and DevOps engineers responsible for running GPU workloads in production. Comfortable with Linux and Kubernetes basics; new to GPU-specific operations.
Prerequisites
- Comfortable with Linux command line and basic shell scripting
- Kubernetes basics (Pods, Deployments, Services), equivalent of CKA prep
- Familiarity with containers (Docker / containerd)
Every lab in this course, module by module
What you break, then what you fix
01Kubernetes Foundations for AI Workloads
AI workloads break Kubernetes assumptions in subtle ways: training pods need predictable scheduling, inference pods need rolling updates without dropped requests, and GPUs are scarce enough that QoS class actually matters. Master the core primitives the way platform engineers do.
- Welcome to NCA-AIIO Labs — Schedule Your First GPU Pod5 min · beginnerSmoke-test your NCA-AIIO lab environment: inspect your isolated Kubernetes cluster, schedule a Pod that requests an NVIDIA GPU, and observe realistic nvidia-smi output. The first 5-minute lab to verify everything works end-to-end.
- Kubernetes Resource Requests & Limits — Who Gets What, and Who Survives40 min · intermediateMaster the most consequential six lines in any Kubernetes manifest: requests, limits, and how they decide scheduling, throttling, eviction, and survival under pressure. Includes the CFS throttling controversy and what 2026 production teams actually do with CPU limits.
- Workload Controllers — Deployment, StatefulSet, DaemonSet for AI35 min · intermediateThree controller types, three workload shapes, three different production failure modes. Learn when to use a Deployment for inference, a StatefulSet for distributed training, and a DaemonSet for per-node GPU infrastructure — and how to spot when someone picked the wrong one.
02The NVIDIA GPU Operator
The GPU Operator is the chain that turns a Helm install into the `nvidia.com/gpu` your workload requests. Inspect every link in that chain (gpu-feature-discovery, device plugin, RuntimeClass, dcgm-exporter), then triage three different broken-pod scenarios where each link breaks differently.
- Inside the NVIDIA GPU Operator — From Helm to Workload-Ready35 min · intermediateWalk the chain that turns a Helm install into the `nvidia.com/gpu` your workload requests. Inspect the cluster the way platform engineers do — node labels, capacity, RuntimeClass — and learn to attribute every piece of evidence to the GPU Operator component that produced it. Finishes with a Triage Day where three broken GPU pods each break a different chain link.
- NVIDIA GPU Operator on k3s: Single-Node Kubernetes for GPU Workloads40 min · intermediateBring up a lightweight single-node Kubernetes cluster with the NVIDIA GPU Operator — k3s install, containerd wiring, Helm values, workload manifests with RBAC and ResourceQuota, plus a full runbook (validation plan, troubleshooting matrix, day-2 ops).
03GPU Scheduling & Sharing
GPUs are scarce and expensive, so placement comes first and contention second. Land workloads on the right hardware across a mixed A100 / L40S fleet with nodeSelector, nodeAffinity and tolerations, resolve GPU contention with PriorityClass and preemption, then share a single card four ways (CUDA streams, multi-process time-slicing, MPS, and MIG).
- Multi-GPU-Type Targeting — nodeSelector, nodeAffinity & Tolerations35 min · intermediateYour fleet has A100s for training, L40S for inference, and Tesla-K80s for dev. Workloads need to land on the right hardware. Master the four primitives Kubernetes gives you — nodeSelector, nodeAffinity (required vs preferred), taints + tolerations — across a real multi-pool cluster.
- PriorityClass & Preemption — Who Survives the GPU Squeeze35 min · intermediateWhen GPU capacity is full and a critical training job lands, who wins? This lab builds the mental model behind PriorityClass and preemption — the only mechanism Kubernetes gives you for resolving GPU contention with intent rather than first-come-first-served. Includes the `preemptionPolicy: Never` escape hatch most teams misuse.
- GPU Sharing: Streams, MPS, MIG, and the Real Cost of Contention45 min · advancedMeasure four ways to share a single GPU — CUDA streams, multi-process time-slicing, MPS, and MIG — and write the production artifacts (start scripts, k8s device-plugin ConfigMaps, MIG geometries) that turn 15%-utilized fleets into 80%-utilized ones.
04Storage & Deployment Strategies
Stop losing your training checkpoints when pods restart, and stop dropping inference requests when you ship a new model. Storage with PVCs and StorageClasses, the three deployment patterns (rolling, blue-green, canary) applied to a real inference Deployment, and a closing triage day where you walk the five-stage pod lifecycle and fix four GPU pods each broken at a different stage.
- Persistent Storage for AI Workloads — PVCs, StorageClass & the Checkpoint Pattern35 min · intermediateStop losing your training checkpoints when pods restart. Learn the PersistentVolumeClaim model end-to-end — StorageClass selection, accessModes (and the RWO-is-per-node trap), the bind/mount/persist lifecycle, and three triage scenarios where storage chains break.
- Rolling Updates, Rollback & Blue-Green for AI Inference35 min · intermediateShip a new model version without dropping requests. Master Kubernetes' three deployment strategies — rolling update with readiness probes, rollback after a bad release, and blue-green via Service-selector swap — all on a stand-in inference Deployment.
- GPU Container Lifecycle: Build, Test, Ship, Rollback40 min · intermediateWalk through the full lifecycle of a production GPU container — multi-stage Dockerfile, self-hosted GPU CI, a fail-fast smoke test, and a Kubernetes Deployment with readiness probes gated on real GPU compute. The pipeline that stops bad images before users see a 500.
- Stuck-Pending Triage Day — Diagnose Any GPU Pod That Won't Run40 min · intermediateThe capstone NCA-AIIO operations lab. Walk the five-stage pod lifecycle (admission → scheduling → image pull → runtime → readiness), learn which `kubectl describe` field signals each stuck point, and finish by fixing four broken GPU pods, each broken at a different stage.
05Model Serving & Inference Capacity
A GPU answering one request at a time is a GPU you are mostly paying to keep warm. This module is about capacity. Build a dynamic batcher by hand so the throughput and latency tradeoff is something you have measured yourself, then run vLLM the way production runs it, with PagedAttention deciding how much KV cache fits on the card, continuous batching keeping the queue full, and the server arguments you hand to Kubernetes. It closes by sweeping batch size against numerical precision on the SKU you actually have, so the number you carry into a capacity plan is your own measurement.
- Inference Serving Patterns: Dynamic Batching, Throughput, and the Triton Mental Model40 min · intermediateBuild a mini-Triton inference server in ~30 lines of Python: a dynamic batcher with max_batch_size and max_queue_delay knobs, load-tested against a naive baseline, swept for the throughput-latency tradeoff, and bridged to a real Triton config.pbtxt.
- vLLM Production Serving: PagedAttention, Continuous Batching, Prefix Caching55 min · advancedStand up vLLM and measure the three features that make it the de-facto inference server: PagedAttention's KV-cache capacity, continuous batching throughput, and prefix caching speedups. Then write the production spec — server args, Kubernetes deployment, monitoring, autoscaling.
- Batch Size & Precision Sweep: Finding Your Sweet Spot40 min · intermediateSweep batch sizes and numerical precisions (fp32, fp16, bf16) on a real model to find the throughput/VRAM knee, then ship a production recommendation with SKU-aware precision picks and an accuracy gate.
06GPU Observability & Performance
Build a real monitoring pipeline (NVML telemetry, dcgm-exporter, Prometheus) and use it to find the actual bottleneck in a training job. Two profilers (PyTorch's built-in and NVIDIA Nsight Systems) cover the full vertical from op-level traces to system-wide kernel timelines.
- GPU Observability: From nvidia-smi to a Production Monitoring Stack40 min · intermediateGo from a raw NVML snapshot to a real monitoring pipeline: capture live GPU telemetry during a workload, diagnose a dataloader bottleneck from the utilization trace, and expose everything as a Prometheus /metrics endpoint.
- GPU Health Checks + Auto-Remediation50 min · advancedBuild a production-grade GPU watchdog: multi-dimensional NVML health probe, rogue-process detection, auto-remediation that kills the offender and verifies recovery, then wire it up with Prometheus alerts and Kubernetes liveness probes.
- Profile PyTorch Training with the Built-in Profiler35 min · intermediateInstrument a training loop with torch.profiler, read the op-level table, inspect the Chrome/Perfetto timeline, and decide when to reach for Nsight Systems instead.
- Nsight Systems Profiling: Finding the Bottleneck That Costs You 40% of Your GPU35 min · intermediateRun the full profile-then-fix loop with NVIDIA Nsight Systems — instrument a training loop with NVTX ranges, capture a .nsys-rep, parse the NVTX summary to pinpoint the bottleneck, then apply a targeted fix and measure the speedup.
07Production Operations & Cost
Run the four-stage cost-audit pipeline that turns raw NVML samples into dollar-denominated waste, and wire MLflow as the experiment tracking and model-registry layer your platform team needs to be the source of truth.
- GPU Cost & Efficiency Audit35 min · intermediateBuild a four-stage cost-audit pipeline — measure, classify, price, recommend — that turns raw NVML samples into dollar-denominated waste and specific remediation actions. The skeleton behind every enterprise GPU cost product.
- MLflow Experiment Tracking: From Single Run to Team Workflow35 min · intermediateWire the four load-bearing pieces of MLflow into a real training loop — tracked runs with params and metrics, a registered model with stage transitions, a multi-run sweep + search, and a production spec (server, k8s Job, tags, autolog).
Lab paths through this course
The same labs, grouped by topic
- Kubernetes GPU course: GPU Operator, scheduling and storage for AIFifteen labs from your first GPU pod to a fleet you can schedule, watch, share and price.
- GPU performance course: CUDA kernels, profiling and sharingSeven GPU labs that find the idle GPU, fix it, go down to the kernel, and price what is left in dollars.
Guides & articles
Deep-dive reading that pairs with this course
AI Infrastructure Engineer Roadmap: Seven Stages from First GPU Pod to Cost Audit
A stage-by-stage AI infrastructure engineer roadmap: Kubernetes for AI workloads, the GPU operator chain, scheduling, storage, serving, observability and cost, with what to build at each stage.
ReadHow to Become an AI Infrastructure Engineer: The GPU and Kubernetes Path
How to become an AI infrastructure engineer from a platform, SRE, or sysadmin background: what the job is, what to learn in order, what to build, and how the interviews actually work.
ReadAI Infrastructure Engineer vs MLOps Engineer: Which Job Are You Applying For?
AI infrastructure engineer vs MLOps engineer by what each one owns when something breaks: the cluster and the GPUs, or the pipeline from experiment to deployed model.
ReadNVIDIA GPU Operator Troubleshooting: Debugging Every Link in the Chain
Why a Kubernetes pod schedules but finds no GPU, why a node shows no nvidia.com/gpu, and how to tell which link in the GPU Operator chain broke, symptom by symptom.
ReadMIG vs MPS vs Time-Slicing: Four Ways to Share a GPU and When Each Wins
CUDA streams, time-slicing, MPS and MIG compared by what they actually give you: concurrency, isolation, or neither. Why naive sharing costs 2x latency for no throughput.
ReadReady to start?
Pro gives you all 21 labs in this path, every other lab on Preporato, and every practice test. $29.99/mo, cancel anytime.