Track · GPU and Kubernetes
Kubernetes GPU projects: GPU Operator, scheduling and CUDA
Your own Kubernetes cluster with GPUs: scheduling, preemption, storage and the GPU Operator, stuck-pod triage and health checks, then CUDA kernels, profiling and GPU sharing on real hardware.
- 20
- Labs
- 12 h
- In total
- Beginner to advanced
- Level
What you will build
- GPU pods scheduled on your own cluster, with requests, limits, priorities and node targeting
- The NVIDIA GPU Operator installed and traced from Helm to a workload-ready node
- Stuck pods diagnosed by the gate they fail at, and automated GPU health checks
- CUDA kernels written and wired into PyTorch
- Bottlenecks found with Nsight and the PyTorch profiler, and a GPU shared with MPS and MIG
Before you start
- A basic Linux shell and the idea of a container image
- What a pod and a node are; the labs introduce every other Kubernetes object as it appears
Tools you will use
KuberneteskubectlHelmNVIDIA GPU Operatork3sCUDANsight SystemsMIG and MPSDALIRAPIDS
Labs in this track
In order, from the first lab to the hardest. Every lab stands on its own, so start wherever you like.
Your GPU cluster
Schedule a first GPU pod, install the GPU Operator, and control who gets which GPU and who is evicted.
- Lab 1Welcome to NCA-AIIO Labs: Schedule Your First GPU PodSmoke-test your NCA-AIIO lab environment: inspect your isolated Kubernetes cluster, schedule a Pod that requests an NVIDIA GPU, and observe realistic nvidia-smi output. The first 5-minute lab to verify everything works end-to-end.5 minBeginnerHostedPro# nca-aiio-welcome · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 2NVIDIA GPU Operator on k3s: Single-Node Kubernetes for GPU WorkloadsBring up a lightweight single-node Kubernetes cluster with the NVIDIA GPU Operator: k3s install, containerd wiring, Helm values, workload manifests with RBAC and ResourceQuota, plus a full runbook (validation plan, troubleshooting matrix, day-2 ops).40 minIntermediateGPUPro
- Lab 3Inside the NVIDIA GPU Operator: From Helm to Workload-ReadyWalk the chain that turns a Helm install into the `nvidia.com/gpu` your workload requests. Inspect the cluster the way platform engineers do: node labels, capacity, RuntimeClass: and learn to attribute every piece of evidence to the GPU Operator component that produced it. Finishes with a Triage Day where three broken GPU pods each break a different chain link.35 minIntermediateHostedPro# nca-aiio-gpu-operator-chain · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 4Kubernetes Resource Requests & Limits: Who Gets What, and Who SurvivesMaster the most consequential six lines in any Kubernetes manifest: requests, limits, and how they decide scheduling, throttling, eviction, and survival under pressure. Includes the CFS throttling controversy and what 2026 production teams actually do with CPU limits.40 minIntermediateHostedPro# nca-aiio-resource-requests-limits · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 5Multi-GPU-Type Targeting: nodeSelector, nodeAffinity & TolerationsYour fleet has A100s for training, L40S for inference, and Tesla-K80s for dev. Workloads need to land on the right hardware. Master the four primitives Kubernetes gives you: nodeSelector, nodeAffinity (required vs preferred), taints + tolerations: across a real multi-pool cluster.35 minIntermediateHostedPro# nca-aiio-multi-gpu-targeting · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 6PriorityClass & Preemption: Who Survives the GPU SqueezeWhen GPU capacity is full and a critical training job lands, who wins? This lab builds the mental model behind PriorityClass and preemption: the only mechanism Kubernetes gives you for resolving GPU contention with intent rather than first-come-first-served. Includes the `preemptionPolicy: Never` escape hatch most teams misuse.35 minIntermediateHostedPro# nca-aiio-priority-preemption · step 1$ lab.check(1)Step 1 Completegrade ........... pass
Run AI workloads
Controllers, storage, containers and rollouts, then triage pods that will not start and keep the GPUs healthy.
- Lab 7Workload Controllers: Deployment, StatefulSet, DaemonSet for AIThree controller types, three workload shapes, three different production failure modes. Learn when to use a Deployment for inference, a StatefulSet for distributed training, and a DaemonSet for per-node GPU infrastructure: and how to spot when someone picked the wrong one.35 minIntermediateHostedPro# nca-aiio-workload-controllers · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 8Persistent Storage for AI Workloads: PVCs, StorageClass & the Checkpoint PatternStop losing your training checkpoints when pods restart. Learn the PersistentVolumeClaim model end-to-end: StorageClass selection, accessModes (and the RWO-is-per-node trap), the bind/mount/persist lifecycle, and three triage scenarios where storage chains break.35 minIntermediateHostedPro
- Lab 9GPU Container Lifecycle: Build, Test, Ship, RollbackWalk through the full lifecycle of a production GPU container: multi-stage Dockerfile, self-hosted GPU CI, a fail-fast smoke test, and a Kubernetes Deployment with readiness probes gated on real GPU compute. The pipeline that stops bad images before users see a 500.40 minIntermediateGPUPro
- Lab 10Rolling Updates, Rollback & Blue-Green for AI InferenceShip a new model version without dropping requests. Master Kubernetes' three deployment strategies: rolling update with readiness probes, rollback after a bad release, and blue-green via Service-selector swap: all on a stand-in inference Deployment.35 minIntermediateHostedPro# nca-aiio-rolling-canary-bluegreen · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 11Stuck-Pending Triage Day: Diagnose Any GPU Pod That Won't RunThe capstone NCA-AIIO operations lab. Walk the five-stage pod lifecycle (admission → scheduling → image pull → runtime → readiness), learn which `kubectl describe` field signals each stuck point, and finish by fixing four broken GPU pods, each broken at a different stage.40 minIntermediateHostedPro# nca-aiio-stuck-pending-triage · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 12GPU Observability: From nvidia-smi to a Production Monitoring StackGo from a raw NVML snapshot to a real monitoring pipeline: capture live GPU telemetry during a workload, diagnose a dataloader bottleneck from the utilization trace, and expose everything as a Prometheus /metrics endpoint.40 minIntermediateGPUProutil86%mem15GBtemp74°Cpower310W
- Lab 13GPU Health Checks + Auto-RemediationBuild a production-grade GPU watchdog: multi-dimensional NVML health probe, rogue-process detection, auto-remediation that kills the offender and verifies recovery, then wire it up with Prometheus alerts and Kubernetes liveness probes.50 minAdvancedGPUPro# NVML · health probe$ pynvml.nvmlDeviceGetPidsrogue_pid ....... 2341VRAM leak: 14.7 GBremediation: SIGKILL → clean
Make the GPU fast
Write CUDA kernels, find the bottleneck with the profiler and Nsight, share a GPU with MPS and MIG, and price the waste.
- Lab 14CUDA Programming FundamentalsWrite four real CUDA C++ kernels and run them from PyTorch: vector add, 2D matrix add, tiled matmul with shared memory, and a custom autograd op.45 minAdvancedGPUPro
- Lab 15Profile PyTorch Training with the Built-in ProfilerInstrument a training loop with torch.profiler, read the op-level table, inspect the Chrome/Perfetto timeline, and decide when to reach for Nsight Systems instead.35 minIntermediateGPUPro# torch.profiler · top ops$ table(sort=cuda_time)gemm ........ 41%memcpy ...... 22%softmax ..... 9%
- Lab 16Nsight Systems Profiling: Finding the Bottleneck That Costs You 40% of Your GPURun the full profile-then-fix loop with NVIDIA Nsight Systems: instrument a training loop with NVTX ranges, capture a .nsys-rep, parse the NVTX summary to pinpoint the bottleneck, then apply a targeted fix and measure the speedup.35 minIntermediateGPUPro
- Lab 17GPU Sharing: Streams, MPS, MIG, and the Real Cost of ContentionMeasure four ways to share a single GPU: CUDA streams, multi-process time-slicing, MPS, and MIG: and write the production artifacts (start scripts, k8s device-plugin ConfigMaps, MIG geometries) that turn 15%-utilized fleets into 80%-utilized ones.45 minAdvancedGPUPro
- Lab 18NVIDIA DALI: GPU-Accelerated Data PipelinesMove image decoding, resizing, and augmentation from CPU to GPU with NVIDIA DALI, and benchmark it against a standard PyTorch DataLoader. The input-pipeline fix that unlocks real multi-GPU throughput.30 minIntermediateGPUPro
- Lab 19GPU-Accelerated Data Science with RAPIDSRewrite a pandas + sklearn data-science pipeline on GPU using cuDF and cuML, benchmark each stage against the CPU baseline, and run an end-to-end filter -> feature-engineer -> predict pipeline that never leaves the GPU.40 minIntermediateGPUPro
- Lab 20GPU Cost & Efficiency AuditBuild a four-stage cost-audit pipeline: measure, classify, price, recommend: that turns raw NVML samples into dollar-denominated waste and specific remediation actions. The skeleton behind every enterprise GPU cost product.35 minIntermediateGPUPro
Graded project
Triage a GPU cluster failing three ways at once
A capture from a shared GPU cluster where three teams are stuck: find each root cause and write the fixes, scored on every rubric criterion.
Guides for this track
Related collections:GPU performance course
Questions about this track
No. Each lab gives you your own cluster or GPU in the browser, ready when the lab opens.
Knowing what a pod and a node are is enough. The welcome lab is a five-minute smoke test, and each lab introduces the objects it uses.
Most take 35 to 45 minutes and save your progress between steps.
GPU scheduling, the GPU Operator, storage and troubleshooting are core to NVIDIA NCA-AIIO, NCP-AIO and NCP-AII.
Other tracks
LLMOps and MLOps
vLLM serving, load tests against SLOs, tracing, drift monitoring and prompt tests in CI.
LLM training and fine-tuning
LoRA and QLoRA fine-tuning, DPO alignment, and a tokenizer and tiny GPT built from scratch.
AI security and red teaming
Prompt injection, tool poisoning and data exfiltration against live targets, then the defenses.
Every lab with Pro
This track and every other one, plus every practice test. $29.99 a month, cancel any time.