Track · LLM training and fine-tuning
How to fine-tune an LLM: hands-on LoRA, QLoRA and DPO projects
Fine-tune with LoRA and QLoRA on real GPUs, align with DPO, prepare and curate the training data, and build a tokenizer and a tiny GPT from scratch to see what training changes.
- 14
- Labs
- 12 h
- In total
- Beginner to advanced
- Level
What you will build
- A Llama model fine-tuned with LoRA and QLoRA, evaluated and merged
- A model aligned with DPO on preference pairs
- A curated, deduplicated instruction dataset and synthetic data you can trust
- A BPE tokenizer and a tiny GPT written from scratch
- A small language model trained from zero on a real GPU
Before you start
- PyTorch basics: tensors, a training loop, what a loss curve is
- The idea of a tokenizer and of train and validation splits
Tools you will use
PyTorchHugging Face TransformersPEFTLoRA and QLoRAbitsandbytesTRLDPOJupyter
Labs in this track
In order, from the first lab to the hardest. Every lab stands on its own, so start wherever you like.
Fine-tune and align
LoRA and QLoRA on a real GPU, quantization, DPO, continued pretraining, an embedding model tuned for retrieval, and the same LoRA idea on images.
- Lab 1Fine-Tune an LLM with LoRA and QLoRA (Jupyter)Fine-tune Meta Llama 3 8B on a custom instruction dataset using LoRA and QLoRA. Learn parameter-efficient fine-tuning from data preparation through evaluation: the #1 most demanded AI skill.45 minIntermediateGPUPro# lora · r=16 · alpha=32$ trainer.train()trainable ....... 0.8%eval_loss ....... 1.24perplexity ...... 3.45
- Lab 2Fine-Tune a Small Text Classifier: Baselines, a Training Loop, Early Stopping, Abstention and a Shift TestTrain a small transformer encoder to label chat messages by intent, and measure what fine-tuning buys: beat a bag-of-words and a frozen-encoder baseline, write the training loop, pick the epoch with the best dev accuracy, send the model's least confident predictions to a person, and test all three models on messages in an unseen style, where the one that learned the meaning pulls ahead.55 minIntermediateHostedPro# finetune-text-classifier · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 3Quantize & Optimize LLMs with bitsandbytesLoad a model in fp16, INT8, and NF4, then benchmark the three precisions on VRAM, latency, and output quality. See where quantization wins and where it costs you.40 minIntermediateGPUPro
- Lab 4RLHF & DPO AlignmentRun real Direct Preference Optimization on a small language model with TRL's DPOTrainer. Capture a baseline, build a preference dataset, train, and measurably shift the model's behavior in four steps.55 minAdvancedGPUPro
- Lab 5Continued Pre-Training: Adapt a Pretrained LM to a New DomainTake GPT-2 and domain-adapt it to Python code in 150 steps, measuring both the gain on code and the cost in catastrophic forgetting on English. The exact recipe behind Code Llama, BloombergGPT, and every domain-specialized LLM of the last three years.45 minAdvancedGPUPro
- Lab 6Train an Embedding Model: Contrastive Fine-Tuning for RetrievalFine-tune a small sentence encoder so it retrieves the right help article from the way people actually ask. Measure recall and MRR on the frozen model, write the in-batch InfoNCE contrastive loss, train the encoder with it, mine the hard negatives it confuses, and keep the checkpoint where dev retrieval peaks, then confirm the gain holds on held-out questions.75 minAdvancedHostedPro# train-embedding-model · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 7Fine-Tune Stable Diffusion with LoRA: Custom Text-to-ImageLoad Stable Diffusion, attach LoRA adapters to the U-Net's attention layers, run a tiny overfit training loop, and generate with the adapted weights to prove that a few million trainable parameters actually move pixels.45 minIntermediateGPUPro
Data for training
The step most tutorials skip: filter, deduplicate, judge, balance and generate the data you train on.
- Lab 8Data Preparation for LLM TrainingBuild a real pretraining/instruction data pipeline: load a raw corpus, apply quality filters, deduplicate, train a BPE tokenizer, and batch-validate on GPU. This is the unglamorous work that actually decides how good your model will be.45 minIntermediateGPUPro# dedup + tokenize$ filter_lang("en")raw ............ 52k → 18ktokenizer ...... BPE 32ktrain_ready .... 16,420
- Lab 9Curate a Fine-Tuning Dataset: Deduplicate, Filter, LLM-Judge and Balance Instruction DataTurn a raw 260-row helpdesk export into a clean instruction-tuning set. Profile it, strip empty, truncated, boilerplate and PII rows with rules, remove near-duplicates with shingles and Jaccard similarity, grade what rules cannot see with an LLM judge calibrated against human labels, then cap categories and write train/eval splits with a data card.50 minIntermediateHostedPro# instruction-data-curation · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 10Synthetic Data Generation for Model TrainingBuild a Self-Instruct style synthetic dataset end-to-end: seed instructions, LLM-driven generation, robust parsing, quality filtering, and dedup + diversity scoring. The same pipeline that produced Alpaca, WizardLM, and most modern instruction-tuning corpora.40 minIntermediateGPUPro
How LLMs work, from scratch
A BPE tokenizer, attention and transformer blocks, a tiny GPT on a CPU and a small model trained on a GPU.
- Lab 11Build a BPE Tokenizer From Scratch in Python (and See Why Some Languages Cost More)Write the byte-level BPE tokenizer behind GPT, Llama and Claude: UTF-8 bytes, pair counting, merges, training with chunking, encoding and decoding. Then measure tokens per word against tiktoken's cl100k and o200k on English and Finnish, and see why the same conversation costs 2.4 times as many tokens in Finnish.45 minBeginnerHostedPro# tokenizer-from-scratch · step 1$ lab.check(1)Step 1 Completegrade ........... pass
- Lab 12Build a Tiny GPT From Scratch on a CPU: Attention, Transformer Blocks and Text GenerationWrite a character-level GPT in PyTorch and train it on the Sherlock Holmes stories in under a minute on one CPU core. Batching for next-character prediction, a bigram baseline, sampling with temperature, a causal self-attention head, multi-head attention, transformer blocks with residuals and LayerNorm, and a measurement of how much each extra character of context lowers the loss.50 minIntermediateHostedPro
- Lab 13Build a Transformer from Scratch: Attention, Masking & LayerNormBuild every piece of a decoder-only transformer by hand: scaled dot-product attention, multi-head attention, the full block with residuals and LayerNorm, then assemble a tiny GPT and train it. No shortcuts, no pre-built attention modules.50 minAdvancedGPUPro
- Lab 14Train a Small Language Model from ScratchTrain a real GPT-style language model from zero on TinyStories: tokenize, wire up the optimizer and LR schedule, run the training loop with validation perplexity, and generate coherent text from your own weights. End-to-end pretraining in minutes on one GPU.55 minAdvancedGPUPro
Guides for this track
Related collections:How to train an LLM from scratch
Questions about this track
The GPU labs start a real NVIDIA GPU for you in the browser. The tokenizer, tiny GPT and data curation labs run on hosted machines.
No. The LoRA lab starts from a base model download and explains each setting as you change it.
Most take 40 to 55 minutes and save your progress between steps.
Fine-tuning, PEFT, alignment and data preparation are core topics of NVIDIA NCP-GENL and NCA-GENL.
Other tracks
Machine learning and deep learning
scikit-learn and PyTorch on real datasets, with the metrics and debugging that decide what ships.
LLMOps and MLOps
vLLM serving, load tests against SLOs, tracing, drift monitoring and prompt tests in CI.
GPU and Kubernetes
GPU scheduling on your own Kubernetes cluster, the GPU Operator, CUDA kernels and profiling.
Every lab with Pro
This track and every other one, plus every practice test. $29.99 a month, cancel any time.