Model Routing & Cost Cascade with NIM
Hands-on lab · Runs in your browser

Model Routing & Cost Cascade with NIM

cost field against an always-large baseline.

Time
25 min
Checked steps
4
Level
Intermediate
Setup
None
ncp-aai
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit25 min · 4 stepsSession running
1 / 4 steps passingHave the model self-rate its confidence · step 2 of 4
main.py▶ Run✓ Check
MID   = "meta/llama-3.3-70b-instruct"LARGE = "nvidia/nemotron-3-super-120b-a12b" def ask(model, question, max_tokens=300):    r = client.chat.completions.create(model=model, messages=[{"role":"user","content":question}], max_tokens=max_tokens, temperature=0)    u = r.usage.model_dump() if hasattr(r.usage,'model_dump') else dict(r.usage or {})    return {"content":(r.choices[0].message.content or "").strip(),"cost":u.get("cost"),"tokens":u.get("total_tokens",0)} def ask_with_confidence(model: str, question: str) -> dict:               
TerminalOutput

4 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Call three tiers and measure cost

    Most production agents spend money on queries the model didn't need to be big for.

  2. 2

    Have the model self-rate its confidence

    To decide *when to escalate*, you need a cheap-to-compute signal that the small model is uncertain.

  3. 3

    Build the cascade

    Now glue the two pieces together.

  4. 4

    Measure savings vs always-large

    The cascade's whole point is cost.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Call three tiers and measure cost

Why cascade?

Most production agents spend money on queries the model didn't need to be big for. "What's 2+2?" runs on an 8B model as well as on a 253B one — but the 253B call costs ~50× more. A cascade flips that economics: try the cheap model first, only pay for the expensive one when the cheap one isn't confident.

NIM exposes several Llama/Nemotron sizes in one OpenAI-compatible catalog. You'll build a 3-tier cascade:

TierModelGood for
Smallmeta/llama-3.1-8b-instructfactual recall, simple classification
Midmeta/llama-3.3-70b-instructmulti-step reasoning, nuance
Largenvidia/nemotron-3-super-120b-a12blong-chain reasoning, edge cases

Each response from NIM includes a usage.cost field — the dollar amount for that call. You'll use it to measure your savings.

Your task

Add to main.py:

  1. Create one OpenAI client pointed at the NIM proxy.
  2. Define a helper ask(model, question, max_tokens=300) that returns a dict {content, cost, tokens}.
  3. Call ask() once per tier on "What is the capital of Japan?".
  4. Print model, content, tokens, and cost for each.

NCP-AAI exam note: Cost-aware routing is part of the Efficiency & Scalability domain. The exam asks about model selection strategy, monitoring cost per call, and tradeoffs between latency, accuracy, and spend.

main.py, the file you edit19 lines
from openai import OpenAI

client = OpenAI(
    base_url="http://nim-proxy.labs.svc:8080/v1",
    api_key="nvapi-inject",
)

SMALL = "meta/llama-3.1-8b-instruct"
MID   = "meta/llama-3.3-70b-instruct"
LARGE = "nvidia/nemotron-3-super-120b-a12b"  # reasoning model, needs extra tokens

# TODO: Define ask(model, question, max_tokens=300) returning {content, cost, tokens}
def ask(model: str, question: str, max_tokens: int = 300) -> dict:
    ______

# Sanity: call all 3 tiers on the same question
for m in [SMALL, MID, LARGE]:
    r = ask(m, "Capital of Japan? One word.")
    print(f"{m:<55}  content={r['content']!r:<35}  tokens={r['tokens']}  cost=${r['cost']}")

Exam domains covered

Efficiency and ScalabilityAgent DevelopmentNVIDIA Platform Implementation

Frequently asked questions

Why use a cascade router instead of always calling the best model?

Because most production queries don't need the best model. On a balanced workload — factual lookups, simple classification, short explanations mixed with occasional hard reasoning — the largest model is overkill on the majority of requests. A 3-tier cascade typically spends 60–80% less on the same traffic because it serves the easy queries from the 8B and the reasoning-heavy queries from the 253B. The cascade never reduces quality on hard queries (those escalate to the large tier anyway); it just stops overpaying on easy ones.

How does confidence-based routing actually work?

You prompt the model to return structured output of the form {"answer": "<short>", "confidence": 1-5} where 5 means "I'm certain" and 1 means "I'd guess." The cascade reads the confidence field and decides whether to accept the tier's answer or escalate. Self-reported confidence isn't perfect — models can be overconfident — but it's well-calibrated enough on factual tasks to route 60–80% of queries correctly, and it costs only one extra field in the response.

Why does the cascade's cost count tiers that didn't answer?

Because they still ran. If the small tier fires, then the mid-tier fires because confidence was low, then the large tier fires and finally answers — the cascade paid for all three calls. A fair comparison against always-large has to charge the cascade for every tier it invoked, not just the winning one. Step 3 tracks cumulative cost across the whole walk, and Step 4's savings percentage is only honest because of that accounting.

What if the small model is overconfident and returns a wrong answer with confidence 5?

That's the cascade's main failure mode. Two mitigations: (a) tune the confidence threshold — in this lab you escalate below 4, but you can make it stricter; (b) add a secondary check like an LLM-as-judge on a sample of small-tier decisions. In production you'd also keep a ground-truth eval set and periodically re-measure small-tier accuracy so overconfidence drift gets caught. The lab's Step 4 dataset deliberately mixes easy and hard questions so you see the overconfidence edge cases.

Which NIM models does this lab cascade across?

Three tiers served through http://nim-proxy.labs.svc:8080/v1: small is meta/llama-3.1-8b-instruct, mid is meta/llama-3.3-70b-instruct, and large is a Llama-3.1 253B-scale tier. They all share the OpenAI-compatible chat completions surface, so switching tiers is a model= string change and nothing else. The per-token pricing table in Step 1 is simplified for the lab but structured so you can replace it with real NIM catalog numbers unchanged.

Does this pattern generalize beyond just two or three model sizes?

Yes. The same logic works for N tiers, for cross-provider cascades (a cheap self-hosted model before a premium API), and for capability-based routing rather than size-based (a code-specialized model before a general one). NeMo Agent Toolkit supports this natively via function groups and workflow-level routing, and many production agent stacks layer cascades on top of an LLM router so cost optimization happens before the agent sees the request. What you build in this lab is the core pattern everything else specializes.

What you'll build in this cost-cascade routing lab

Model routing and cost cascades are the single highest-leverage optimisation on most production LLM apps — real teams cut inference spend 60–80% on balanced workloads by serving the easy queries from a small model and reserving the expensive tier for genuinely hard cases. This lab builds a three-tier confidence-aware cascade against NVIDIA NIM endpoints we provision, measures real token-based cost per tier, and benchmarks the cascade against an always-large baseline on a mixed-difficulty question set. You walk away with a working cascade(question) function, cost numbers you can quote, and the mental model for when confidence-based routing wins versus always-use-the-best-model.

The substance is confidence-driven routing. The small tier — meta/llama-3.1-8b-instruct — handles factual recall and classification. The mid tier — meta/llama-3.3-70b-instruct — absorbs moderate reasoning. The large tier — Llama-3.1 253B-scale — is the reserve for long-chain reasoning and edge cases. You prompt each tier to return structured JSON like {"answer": "<short>", "confidence": 1-5}, walk small → mid → large stopping when confidence clears a threshold, and — critically — charge the cascade for every tier it invoked, not just the winning one. You'll see why self-reported confidence is surprisingly well-calibrated on factual tasks, why the main failure mode is overconfident wrong answers (and how LLM-as-judge on a sample catches them), and why this primitive generalises to cross-provider cascades, capability-based routing, and NeMo Agent Toolkit function-group workflows.

Prerequisites: Python, comfort with JSON payloads, a basic feel for token-based LLM pricing, and prior NIM exposure (the react-agent-nim lab works). The hosted environment ships with the OpenAI Python SDK pointed at our managed NIM proxy — all three tiers share the same OpenAI-compatible endpoint, so swapping between them is a model= string change. No GPU provisioning. About 25 minutes of focused work. You leave with per-tier dollar costs on real token counts, a clean structured-output confidence signal, a cascade that short-circuits on high confidence, and a side-by-side savings report against always-large — the shape of number you walk into a finance review with.