Step 1: Call three tiers and measure cost
Why cascade?
Most production agents spend money on queries the model didn't need to be big for. "What's 2+2?" runs on an 8B model as well as on a 253B one — but the 253B call costs ~50× more. A cascade flips that economics: try the cheap model first, only pay for the expensive one when the cheap one isn't confident.
NIM exposes several Llama/Nemotron sizes in one OpenAI-compatible catalog. You'll build a 3-tier cascade:
| Tier | Model | Good for |
|---|---|---|
| Small | meta/llama-3.1-8b-instruct | factual recall, simple classification |
| Mid | meta/llama-3.3-70b-instruct | multi-step reasoning, nuance |
| Large | nvidia/nemotron-3-super-120b-a12b | long-chain reasoning, edge cases |
Each response from NIM includes a usage.cost field — the dollar amount for that call. You'll use it to measure your savings.
Your task
Add to main.py:
- Create one OpenAI client pointed at the NIM proxy.
- Define a helper
ask(model, question, max_tokens=300)that returns a dict{content, cost, tokens}. - Call
ask()once per tier on"What is the capital of Japan?". - Print model, content, tokens, and cost for each.
NCP-AAI exam note: Cost-aware routing is part of the Efficiency & Scalability domain. The exam asks about model selection strategy, monitoring cost per call, and tradeoffs between latency, accuracy, and spend.
main.py, the file you edit19 lines
from openai import OpenAI
client = OpenAI(
base_url="http://nim-proxy.labs.svc:8080/v1",
api_key="nvapi-inject",
)
SMALL = "meta/llama-3.1-8b-instruct"
MID = "meta/llama-3.3-70b-instruct"
LARGE = "nvidia/nemotron-3-super-120b-a12b" # reasoning model, needs extra tokens
# TODO: Define ask(model, question, max_tokens=300) returning {content, cost, tokens}
def ask(model: str, question: str, max_tokens: int = 300) -> dict:
______
# Sanity: call all 3 tiers on the same question
for m in [SMALL, MID, LARGE]:
r = ask(m, "Capital of Japan? One word.")
print(f"{m:<55} content={r['content']!r:<35} tokens={r['tokens']} cost=${r['cost']}")