LLM inference optimization course: vLLM, quantization and profiling
Nine labs on the numbers that decide whether an LLM deployment is affordable: VRAM, latency, throughput, dollars.
Part of the AI Engineer Course.
What you'll build
- fp16 vs INT8 vs NF4 benchmarked on VRAM, latency and quality with bitsandbytes
- The batch-size / precision knee for a real model, with a production recommendation
- A mini dynamic batcher (the Triton mental model) built and load-tested in ~30 lines
- vLLM in production: PagedAttention capacity, continuous batching, prefix caching
- Profiles from torch.profiler and Nsight Systems, and a GPU cost audit in dollars
About this path
Serving a language model is an optimization problem: how many requests per second can one GPU answer, at what latency, at what memory footprint, and at what cost per thousand tokens. Quantization (storing weights in 8-bit or 4-bit instead of 16-bit) shrinks memory and often speeds decoding; batching and paged key-value caches raise throughput; profilers show where the GPU is idle. These are the levers the NVIDIA NCP-GENL Model Optimization domain tests, and the ones our practice-test data shows candidates miss most.
The labs measure everything rather than asserting it. Load a model in fp16, INT8 and NF4 and benchmark VRAM, latency and output quality; sweep batch size and precision to find the throughput knee; build a mini Triton-style dynamic batcher in about thirty lines and load-test it; move to production serving with vLLM and measure PagedAttention capacity, continuous batching and prefix caching; cascade requests across model tiers on NIM and read the real cost field. Three labs close the loop on efficiency: the PyTorch profiler, Nsight Systems, and a four-stage GPU cost audit that turns NVML samples into dollars of waste.
GPU labs run on live NVIDIA pods in the browser; the routing lab is hosted. Each takes 25 to 55 minutes and is checked step by step.
Who should join
- Python and basic PyTorch
- What a token, a batch and a KV cache are (the labs define them again in context)
- No serving experience required
Every step is checked against the live environment. Progress saves between sessions.
Outline
9 labs · about 6 hours
1. Shrink and measure
Quantize, sweep precision and batch size, and read the numbers that decide deployment.
2. Serve
From a hand-built dynamic batcher to vLLM, plus cost routing across model tiers.
- 3Inference Serving Patterns: Dynamic Batching, Throughput, and the Triton Mental ModelGPU lab · 40 min · IntermediatePro
- 4Deploy & Serve LLMs in Production (Jupyter)GPU lab · 45 min · IntermediatePro
- 5vLLM Production Serving: PagedAttention, Continuous Batching, Prefix CachingGPU lab · 55 min · AdvancedPro
- 6Model Routing & Cost Cascade with NIMHosted lab · 25 min · IntermediatePro
3. Profile and pay less
Find the idle GPU with the profiler and Nsight, then turn NVML samples into dollars of waste.
Run all 9 labs with Preporato Pro, plus every other lab and practice test.
$29.99 per month or $290 per year. Cancel any time.
Frequently asked questions
Yes. GPU labs provision a dedicated NVIDIA GPU pod for your session, so the VRAM, latency and throughput numbers you measure are real, not simulated. The routing lab is hosted and calls NVIDIA NIM.
Model Optimization and GPU Acceleration are the two heaviest NCP-GENL domains, and quantization, batching, serving patterns and profiling appear across NCP-GENL, NCA-GENL and the NVIDIA infrastructure certifications.
The serving labs use vLLM, a hand-built dynamic batcher, and NVIDIA NIM; the profiling labs use torch.profiler and Nsight Systems. TensorRT-LLM concepts are covered in the exam guides on this site rather than as a lab.
The labs are included in Preporato Pro ($29.99 per month or $290 per year) with every practice test and lab on the platform.