LLM inference optimization course: vLLM, quantization and profiling

Nine labs on the numbers that decide whether an LLM deployment is affordable: VRAM, latency, throughput, dollars.

Part of the AI Engineer Course.

What you'll build

  • fp16 vs INT8 vs NF4 benchmarked on VRAM, latency and quality with bitsandbytes
  • The batch-size / precision knee for a real model, with a production recommendation
  • A mini dynamic batcher (the Triton mental model) built and load-tested in ~30 lines
  • vLLM in production: PagedAttention capacity, continuous batching, prefix caching
  • Profiles from torch.profiler and Nsight Systems, and a GPU cost audit in dollars

About this path

Serving a language model is an optimization problem: how many requests per second can one GPU answer, at what latency, at what memory footprint, and at what cost per thousand tokens. Quantization (storing weights in 8-bit or 4-bit instead of 16-bit) shrinks memory and often speeds decoding; batching and paged key-value caches raise throughput; profilers show where the GPU is idle. These are the levers the NVIDIA NCP-GENL Model Optimization domain tests, and the ones our practice-test data shows candidates miss most.

The labs measure everything rather than asserting it. Load a model in fp16, INT8 and NF4 and benchmark VRAM, latency and output quality; sweep batch size and precision to find the throughput knee; build a mini Triton-style dynamic batcher in about thirty lines and load-test it; move to production serving with vLLM and measure PagedAttention capacity, continuous batching and prefix caching; cascade requests across model tiers on NIM and read the real cost field. Three labs close the loop on efficiency: the PyTorch profiler, Nsight Systems, and a four-stage GPU cost audit that turns NVML samples into dollars of waste.

GPU labs run on live NVIDIA pods in the browser; the routing lab is hosted. Each takes 25 to 55 minutes and is checked step by step.

Who should join

  • Python and basic PyTorch
  • What a token, a batch and a KV cache are (the labs define them again in context)
  • No serving experience required

Every step is checked against the live environment. Progress saves between sessions.

Outline

9 labs · about 6 hours

Run all 9 labs with Preporato Pro, plus every other lab and practice test.

$29.99 per month or $290 per year. Cancel any time.

Frequently asked questions