Track · RAG and search
RAG pipeline projects: build, evaluate and ship RAG
From your first vector search to a retriever you can measure and ship: chunking, hybrid search with reranking, GraphRAG, permission-aware retrieval and an evaluation gate that catches regressions.
- 14
- Labs
- 12 h
- In total
- Intermediate to advanced
- Level
- 1
- Free
What you will build
- A working RAG pipeline on hosted NVIDIA NIM, with an agent answering from your knowledge base
- A chunking and parsing setup chosen by measured recall on real PDFs
- A hybrid retriever: BM25 plus dense search, fused with RRF and reranked
- GraphRAG answers to multi-hop questions that vector search misses
- An evaluation harness with a regression gate, and retrieval that respects document permissions
Before you start
- Python basics and comfort reading a short script
- What an embedding is: a vector that places similar text near similar text
Tools you will use
NVIDIA NIMBGE embeddingsBM25Reciprocal Rank FusionCross-encodersHyDEHNSW and IVFGraphRAGNeMo Retrieverpypdf
Labs in this track
In order, from the first lab to the hardest. Every lab stands on its own, so start wherever you like.
Build the pipeline
Stand up RAG end to end, first on hosted NVIDIA NIM (free), then with local models on a GPU, and learn what chunking and parsing do to recall.
- Lab 1Build a RAG Pipeline with NVIDIA NIMBuild a complete Retrieval Augmented Generation pipeline: from document chunking to vector search to an agent that answers questions from your knowledge base.35 minIntermediateHostedFree
- Lab 2Retrieval-Augmented Generation (RAG) Pipeline with Local ModelsBuild an end-to-end RAG pipeline on a single GPU: BGE embeddings, L2-normalized vector retrieval by dot product, and a local generator that answers with and without retrieved context so you can see exactly what retrieval changes.45 minIntermediateGPUPro
- Lab 3RAG Chunking Shoot-Out: Measure Which Chunking Strategy RetrievesMeasure chunking strategies for retrieval-augmented generation instead of guessing: write a recall metric over questions with exact evidence sentences, then compare whole documents, fixed-size chunks with and without overlap, heading-aware section chunks and sections with their heading path, on recall and on characters sent, and prove your pick on questions it has never seen.50 minIntermediateHostedPro
- Lab 4Document AI: Turn PDFs, Tables and Scans into Clean RAG ChunksBuild a document ingestion pipeline with pypdf and a vision model: extract page text, strip running headers and footers, repair hyphenated line breaks, turn table rows into self-describing records, transcribe a scanned page, and produce chunks whose evidence can be found and retrieved.55 minIntermediateHostedPro
Make retrieval good
Hybrid search, reranking, query rewriting, index trade-offs, graphs and images: the techniques that decide whether the right chunk comes back.
- Lab 5Hybrid Search and Reranking: BM25, Embeddings, RRF and an LLM RerankerBuild hybrid retrieval step by step and measure every stage: BM25 keyword search, dense embedding search, reciprocal rank fusion, and a listwise LLM reranker that orders the candidate pool. Size the rerank pool from a recall and prompt-length report and prove the pipeline on questions it has never seen.55 minIntermediateHostedPro
- Lab 6Advanced RAG: Hybrid Search + Cross-Encoder RerankingBuild a production-shape retrieval stack: dense bi-encoder plus from-scratch BM25, fused with Reciprocal Rank Fusion, then re-ordered by a BAAI cross-encoder. The exact architecture behind modern enterprise RAG.40 minAdvancedGPUPro
- Lab 7Query Rewriting and HyDE: Better Queries for the Search You HaveImprove retrieval without touching the index: rewrite follow-up questions into standalone ones, search several model-written phrasings and fuse them, search with a hypothetical answer (HyDE), then fold rewriting and HyDE into one model call per question and prove it on unseen questions.50 minIntermediateHostedPro
Lab 8Vector Index Internals: Exact Search, IVF, HNSW, Product Quantization and Filtered Search with FAISSMeasure what each vector index trades for speed on 100,000 product embeddings. Build exact search and recall@10, sweep IVF nprobe and HNSW efSearch for recall against latency, compress with product quantization and win recall back by re-ranking from disk, and filter to in-stock products inside the search.55 minIntermediateHostedPro
Lab 9GraphRAG: Build a Knowledge Graph from a Wiki with an LLM and Answer Multi-Hop QuestionsTurn an engineering wiki into a knowledge graph and answer the questions vector RAG cannot. Extract triples with an LLM, resolve names against a service catalogue, score the graph against hand labels, walk it with transitive hops, have the model translate questions into graph queries, and compare against vector RAG by question type.60 minIntermediateHostedPro
Lab 10Multimodal RAG with NeMo RetrieverBuild an image-query RAG system: embed a catalog with NeMo Retriever, translate an uploaded image into a retrieval query via a VLM, and ground the VLM's final answer in the retrieved passages.35 minIntermediateHostedPro
Lab 11Embedding Recommender: More-Like-This, Offline Evaluation, Co-occurrence, Cold Start and MMRBuild a bookshop's recommendations from blurb embeddings and click history: more-like-this by cosine similarity, hit rate and MRR on held-out sessions, taste profiles, co-occurrence, a count-weighted blend that still recommends books nobody has clicked, and MMR and author caps for lists people want.50 minIntermediateHostedPro
Ship it
Measure it, gate it and lock it down: an evaluation harness, permission-aware retrieval and the long-context question.
Lab 12Evaluate a RAG Pipeline: Retrieval Metrics, Faithfulness and a Regression GateTurn 'the RAG bot seems better' into numbers you can gate on: chunk a knowledge base with paragraph provenance, label a golden set, score retrieval with recall, precision, MRR and nDCG, sweep chunk sizes, judge answer faithfulness claim by claim with a model, grade correctness, triage every failure by where it happened, and block a candidate configuration that regresses a protected question.80 minIntermediateHostedPro
Lab 13Permission-Aware RAG: Access Control, Metadata Filters and a Cache That Cannot LeakEnforce document permissions in a retrieval-augmented assistant: write allow, deny and clearance rules, filter by region and effective dates, search only what each user may see, answer with citations and an exact refusal, run leak probes, and build an answer cache keyed so it can never serve one user's answer to another.50 minIntermediateHostedPro
Lab 14Long Context vs RAG: Accuracy, Cost and Latency on the Same 24 QuestionsAnswer the same questions from a 19,000-token handbook three ways and measure each. Put the whole document in the prompt, then retrieve sections, attach amendments to the sections they change, and handle list questions with a map step over a few sections at a time. Grade every answer, price it and time it.55 minIntermediateHostedPro
Attack and defend RAG
Three labs from the AI security track break into a RAG assistant like the one you built, then close the holes.
Graded project
Build a RAG pipeline on your own
Ingest and chunk the documents, retrieve, and answer three questions with citations. Every criterion is scored with evidence from your code.
Guides for this track
Related collections:NVIDIA NIM vs NeMo
Questions about this track
No. GPU labs run on an NVIDIA GPU started for you in the browser, and hosted labs call the models through the platform. There is nothing to install and no keys to manage.
The hosted RAG build with NVIDIA NIM is free. The other labs in this track are part of Pro, which also opens every other lab and practice test.
Most labs take 35 to 60 minutes and save your progress, so you can stop between steps and pick up where you left off.
RAG design, chunking, retrieval evaluation and grounding appear in the NVIDIA NCP-AAI and NCP-GENL blueprints and in the Integration domain of Anthropic's CCAR-P.
Other tracks
AI agents and MCP
Tool loops, ReAct, MCP servers and clients, multi-agent supervisors and agent evaluation.
LLMOps and MLOps
vLLM serving, load tests against SLOs, tracing, drift monitoring and prompt tests in CI.
AI security and red teaming
Prompt injection, tool poisoning and data exfiltration against live targets, then the defenses.
Every lab with Pro
This track and every other one, plus every practice test. $29.99 a month, cancel any time.