GPU performance course: CUDA kernels, profiling and sharing

Seven GPU labs that find the idle GPU, fix it, go down to the kernel, and price what is left in dollars.

Part of the AI Infrastructure Engineer Course.

What you'll build

  • Four CUDA C++ kernels (vector add, 2D matrix add, tiled shared-memory matmul, custom autograd op) called from PyTorch
  • A torch.profiler op table and Perfetto timeline for a training loop, and an Nsight Systems .nsys-rep with the bottleneck found, fixed and the speedup measured
  • The batch-size / precision knee for a real model, with SKU-aware precision picks and an accuracy gate
  • A DALI GPU input pipeline benchmarked against the standard PyTorch DataLoader
  • Streams, time-slicing, MPS and MIG measured on one GPU with production ConfigMaps and MIG geometries, and a cost audit that prices GPU waste in dollars

About this path

GPU performance work comes down to one question: is the GPU busy, and if it is idle, why? Answering it means knowing what a CUDA kernel is (a function that runs in parallel across thousands of GPU threads), how to read a profiler timeline, how batch size and numeric precision (fp32, fp16, bf16) move throughput and memory, and how several jobs can share one device through CUDA streams, MPS (Multi-Process Service, which lets several processes run kernels on one GPU concurrently) and MIG (Multi-Instance GPU, which partitions one GPU into isolated slices). Those are the topics behind the NVIDIA NCP-GENL GPU Acceleration domain (performance profiling, memory optimization, Tensor Core utilization) and the MIG, GPU monitoring and troubleshooting items on the NCP-AII and NCA-AIIO infrastructure exams. The labs make each of those concrete: you capture the trace that shows the idle gap and apply the fix that closes it.

The collection measures first and fixes second. Instrument a training loop with torch.profiler, read the op-level table and the Chrome/Perfetto timeline, and learn when to reach for Nsight Systems; then run the full profile-then-fix loop in Nsight Systems: mark the loop with NVTX ranges (named regions that show up in the trace), capture a .nsys-rep, pinpoint the bottleneck, apply a targeted fix and measure the speedup. Two labs then tune the input side and the numeric side: move image decoding, resizing and augmentation onto the GPU with NVIDIA DALI and benchmark it against a standard PyTorch DataLoader; sweep batch size and fp32/fp16/bf16 precision on a real model to find the throughput/VRAM knee and ship a recommendation with SKU-aware precision picks and an accuracy gate. The two advanced labs go lower and wider: write four CUDA C++ kernels (vector add, 2D matrix add, tiled matmul with shared memory, a custom autograd op) and run them from PyTorch; then measure four ways to share one GPU (CUDA streams, time-slicing, MPS, MIG) and write the start scripts, Kubernetes device-plugin ConfigMaps and MIG geometries that put sharing into production. The last lab builds a four-stage cost audit (measure, classify, price, recommend) that turns raw NVML samples into dollar-denominated waste and remediation actions.

All seven labs run on dedicated NVIDIA GPU pods in the browser, so every profile and benchmark comes from a real device; each takes 30 to 45 minutes and is checked step by step.

Who should join

  • Python and basic PyTorch (a training loop you can read)
  • For the CUDA lab, comfort reading C or C++; the four kernels are short
  • No profiling or GPU-operations experience required; each tool is introduced inside its lab

Every step is checked against the live environment. Progress saves between sessions.

Outline

7 labs · about 4 hours

1. Find where the GPU is idle

Read the torch.profiler table and timeline, then run the full Nsight Systems profile-then-fix loop and measure the speedup.

  1. 1Profile PyTorch Training with the Built-in ProfilerGPU lab · 35 min · IntermediatePro
  2. 2Nsight Systems Profiling: Finding the Bottleneck That Costs You 40% of Your GPUGPU lab · 35 min · IntermediatePro

2. Feed it and size it

Move the input pipeline onto the GPU with DALI, then sweep batch size and precision to the throughput/VRAM knee.

  1. 3NVIDIA DALI: GPU-Accelerated Data PipelinesGPU lab · 30 min · IntermediatePro
  2. 4Batch Size & Precision Sweep: Finding Your Sweet SpotGPU lab · 40 min · IntermediatePro

3. Go down to the kernel, then share the device

Write real CUDA kernels and call them from PyTorch, then measure streams, time-slicing, MPS and MIG on one GPU and write the production artifacts.

  1. 5CUDA Programming FundamentalsGPU lab · 45 min · AdvancedPro
  2. 6GPU Sharing: Streams, MPS, MIG, and the Real Cost of ContentionGPU lab · 45 min · AdvancedPro

4. Turn utilization into dollars

A four-stage audit that turns NVML samples into dollar-denominated waste and specific remediation actions.

  1. 7GPU Cost & Efficiency AuditGPU lab · 35 min · IntermediatePro

Run all 7 labs with Preporato Pro, plus every other lab and practice test.

$29.99 per month or $290 per year. Cancel any time.

Frequently asked questions