Track · GPU and Kubernetes

Kubernetes GPU projects: GPU Operator, scheduling and CUDA

Your own Kubernetes cluster with GPUs: scheduling, preemption, storage and the GPU Operator, stuck-pod triage and health checks, then CUDA kernels, profiling and GPU sharing on real hardware.

20
Labs
12 h
In total
Beginner to advanced
Level
Open lab 1Welcome to NCA-AIIO Labs: Schedule Your First GPU Pod · 5 min

What you will build

  • GPU pods scheduled on your own cluster, with requests, limits, priorities and node targeting
  • The NVIDIA GPU Operator installed and traced from Helm to a workload-ready node
  • Stuck pods diagnosed by the gate they fail at, and automated GPU health checks
  • CUDA kernels written and wired into PyTorch
  • Bottlenecks found with Nsight and the PyTorch profiler, and a GPU shared with MPS and MIG

Before you start

  • A basic Linux shell and the idea of a container image
  • What a pod and a node are; the labs introduce every other Kubernetes object as it appears

Tools you will use

KuberneteskubectlHelmNVIDIA GPU Operatork3sCUDANsight SystemsMIG and MPSDALIRAPIDS

Labs in this track

In order, from the first lab to the hardest. Every lab stands on its own, so start wherever you like.

1

Your GPU cluster

Schedule a first GPU pod, install the GPU Operator, and control who gets which GPU and who is evicted.

  1. # nca-aiio-welcome · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 1Welcome to NCA-AIIO Labs: Schedule Your First GPU PodSmoke-test your NCA-AIIO lab environment: inspect your isolated Kubernetes cluster, schedule a Pod that requests an NVIDIA GPU, and observe realistic nvidia-smi output. The first 5-minute lab to verify everything works end-to-end.5 minBeginnerHostedPro
  2. k3sop
    Lab 2NVIDIA GPU Operator on k3s: Single-Node Kubernetes for GPU WorkloadsBring up a lightweight single-node Kubernetes cluster with the NVIDIA GPU Operator: k3s install, containerd wiring, Helm values, workload manifests with RBAC and ResourceQuota, plus a full runbook (validation plan, troubleshooting matrix, day-2 ops).40 minIntermediateGPUPro
  3. # nca-aiio-gpu-operator-chain · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 3Inside the NVIDIA GPU Operator: From Helm to Workload-ReadyWalk the chain that turns a Helm install into the `nvidia.com/gpu` your workload requests. Inspect the cluster the way platform engineers do: node labels, capacity, RuntimeClass: and learn to attribute every piece of evidence to the GPU Operator component that produced it. Finishes with a Triage Day where three broken GPU pods each break a different chain link.35 minIntermediateHostedPro
  4. # nca-aiio-resource-requests-limits · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 4Kubernetes Resource Requests & Limits: Who Gets What, and Who SurvivesMaster the most consequential six lines in any Kubernetes manifest: requests, limits, and how they decide scheduling, throttling, eviction, and survival under pressure. Includes the CFS throttling controversy and what 2026 production teams actually do with CPU limits.40 minIntermediateHostedPro
  5. # nca-aiio-multi-gpu-targeting · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 5Multi-GPU-Type Targeting: nodeSelector, nodeAffinity & TolerationsYour fleet has A100s for training, L40S for inference, and Tesla-K80s for dev. Workloads need to land on the right hardware. Master the four primitives Kubernetes gives you: nodeSelector, nodeAffinity (required vs preferred), taints + tolerations: across a real multi-pool cluster.35 minIntermediateHostedPro
  6. # nca-aiio-priority-preemption · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 6PriorityClass & Preemption: Who Survives the GPU SqueezeWhen GPU capacity is full and a critical training job lands, who wins? This lab builds the mental model behind PriorityClass and preemption: the only mechanism Kubernetes gives you for resolving GPU contention with intent rather than first-come-first-served. Includes the `preemptionPolicy: Never` escape hatch most teams misuse.35 minIntermediateHostedPro
2

Run AI workloads

Controllers, storage, containers and rollouts, then triage pods that will not start and keep the GPUs healthy.

  1. # nca-aiio-workload-controllers · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 7Workload Controllers: Deployment, StatefulSet, DaemonSet for AIThree controller types, three workload shapes, three different production failure modes. Learn when to use a Deployment for inference, a StatefulSet for distributed training, and a DaemonSet for per-node GPU infrastructure: and how to spot when someone picked the wrong one.35 minIntermediateHostedPro
  2. Lab 8Persistent Storage for AI Workloads: PVCs, StorageClass & the Checkpoint PatternStop losing your training checkpoints when pods restart. Learn the PersistentVolumeClaim model end-to-end: StorageClass selection, accessModes (and the RWO-is-per-node trap), the bind/mount/persist lifecycle, and three triage scenarios where storage chains break.35 minIntermediateHostedPro
  3. buildtestshiproll
    Lab 9GPU Container Lifecycle: Build, Test, Ship, RollbackWalk through the full lifecycle of a production GPU container: multi-stage Dockerfile, self-hosted GPU CI, a fail-fast smoke test, and a Kubernetes Deployment with readiness probes gated on real GPU compute. The pipeline that stops bad images before users see a 500.40 minIntermediateGPUPro
  4. # nca-aiio-rolling-canary-bluegreen · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 10Rolling Updates, Rollback & Blue-Green for AI InferenceShip a new model version without dropping requests. Master Kubernetes' three deployment strategies: rolling update with readiness probes, rollback after a bad release, and blue-green via Service-selector swap: all on a stand-in inference Deployment.35 minIntermediateHostedPro
  5. # nca-aiio-stuck-pending-triage · step 1
    $ lab.check(1)
    Step 1 Complete
    grade ........... pass
    Lab 11Stuck-Pending Triage Day: Diagnose Any GPU Pod That Won't RunThe capstone NCA-AIIO operations lab. Walk the five-stage pod lifecycle (admission → scheduling → image pull → runtime → readiness), learn which `kubectl describe` field signals each stuck point, and finish by fixing four broken GPU pods, each broken at a different stage.40 minIntermediateHostedPro
  6. util
    86%
    mem
    15GB
    temp
    74°C
    power
    310W
    Lab 12GPU Observability: From nvidia-smi to a Production Monitoring StackGo from a raw NVML snapshot to a real monitoring pipeline: capture live GPU telemetry during a workload, diagnose a dataloader bottleneck from the utilization trace, and expose everything as a Prometheus /metrics endpoint.40 minIntermediateGPUPro
  7. # NVML · health probe
    $ pynvml.nvmlDeviceGetPids
    rogue_pid ....... 2341
    VRAM leak: 14.7 GB
    remediation: SIGKILL → clean
    Lab 13GPU Health Checks + Auto-RemediationBuild a production-grade GPU watchdog: multi-dimensional NVML health probe, rogue-process detection, auto-remediation that kills the offender and verifies recovery, then wire it up with Prometheus alerts and Kubernetes liveness probes.50 minAdvancedGPUPro
3

Make the GPU fast

Write CUDA kernels, find the bottleneck with the profiler and Nsight, share a GPU with MPS and MIG, and price the waste.

  1. naiveaddtiledautograd
    Lab 14CUDA Programming FundamentalsWrite four real CUDA C++ kernels and run them from PyTorch: vector add, 2D matrix add, tiled matmul with shared memory, and a custom autograd op.45 minAdvancedGPUPro
  2. # torch.profiler · top ops
    $ table(sort=cuda_time)
    gemm ........ 41%
    memcpy ...... 22%
    softmax ..... 9%
    Lab 15Profile PyTorch Training with the Built-in ProfilerInstrument a training loop with torch.profiler, read the op-level table, inspect the Chrome/Perfetto timeline, and decide when to reach for Nsight Systems instead.35 minIntermediateGPUPro
  3. Lab 16Nsight Systems Profiling: Finding the Bottleneck That Costs You 40% of Your GPURun the full profile-then-fix loop with NVIDIA Nsight Systems: instrument a training loop with NVTX ranges, capture a .nsys-rep, parse the NVTX summary to pinpoint the bottleneck, then apply a targeted fix and measure the speedup.35 minIntermediateGPUPro
  4. GPUABC
    Lab 17GPU Sharing: Streams, MPS, MIG, and the Real Cost of ContentionMeasure four ways to share a single GPU: CUDA streams, multi-process time-slicing, MPS, and MIG: and write the production artifacts (start scripts, k8s device-plugin ConfigMaps, MIG geometries) that turn 15%-utilized fleets into 80%-utilized ones.45 minAdvancedGPUPro
  5. CPUDALI
    Lab 18NVIDIA DALI: GPU-Accelerated Data PipelinesMove image decoding, resizing, and augmentation from CPU to GPU with NVIDIA DALI, and benchmark it against a standard PyTorch DataLoader. The input-pipeline fix that unlocks real multi-GPU throughput.30 minIntermediateGPUPro
  6. pandascuDF
    Lab 19GPU-Accelerated Data Science with RAPIDSRewrite a pandas + sklearn data-science pipeline on GPU using cuDF and cuML, benchmark each stage against the CPU baseline, and run an end-to-end filter -> feature-engineer -> predict pipeline that never leaves the GPU.40 minIntermediateGPUPro
  7. $$$
    Lab 20GPU Cost & Efficiency AuditBuild a four-stage cost-audit pipeline: measure, classify, price, recommend: that turns raw NVML samples into dollar-denominated waste and specific remediation actions. The skeleton behind every enterprise GPU cost product.35 minIntermediateGPUPro

Graded project

Triage a GPU cluster failing three ways at once

A capture from a shared GPU cluster where three teams are stuck: find each root cause and write the fixes, scored on every rubric criterion.

Open the project

Guides for this track

Related collections:GPU performance course

Questions about this track

Other tracks

Every lab with Pro

This track and every other one, plus every practice test. $29.99 a month, cancel any time.

See Pro