Visual Q&A with NVIDIA VLMs
Hands-on lab · Runs in your browser

Visual Q&A with NVIDIA VLMs

Send images to a Vision-Language Model via NIM, answer questions about them, extract structured fields from a receipt-style image, and compare two VLMs on the same task — all through the OpenAI-compatible chat endpoint.

Time
30 min
Checked steps
4
Level
Intermediate
Setup
None
ncp-aainca-genm
Read step 1

Hands-on labs require Pro · $29.99/mo · cancel anytime

Lab cockpit30 min · 4 stepsSession running
0 / 4 steps passingSend one image and ask a question · step 1 of 4
main.py▶ Run✓ Check
make_image()data_url = image_to_data_url("/tmp/test_img.jpg") client = OpenAI(base_url="http://nim-proxy.labs.svc:8080/v1", api_key="nvapi-inject")VLM_MODEL = "google/gemma-4-31b-it" def describe_image(prompt: str, data_url: str) -> str:           
TerminalOutput

4 steps, each checked when you finish it

A check runs your work at the end of every step. Hints and the full solution are there if you get stuck.

  1. 1

    Send one image and ask a question

    Vision-Language Models use the same OpenAI-compatible chat endpoint as text-only models, but the content field inside a user message becomes a list of content parts:

  2. 2

    Reason across two images

    A VLM can compare images in a single turn — you simply pass more than one image_url part in the content list.

  3. 3

    Structured extraction from an image

    VLMs can produce structured output via the same tools / tool_choice primitives you used in structured-output-tools — the schema is enforced at the API boundary, so you skip the regex-and-pray parsing dance.

  4. 4

    Compare two VLMs with different extraction strategies

    Not every VLM supports the tools / tool_choice API.

Step 1 as it appears in the lab

The lab’s own text. The hint and the solution stay inside the lab.

Step 1: Send one image and ask a question

Calling a VLM

Vision-Language Models use the same OpenAI-compatible chat endpoint as text-only models, but the content field inside a user message becomes a list of content parts:

{
  "role": "user",
  "content": [
    {"type": "text", "text": "What's in this image?"},
    {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,<...>"}},
  ]
}

You can pass an image as a data URL (base64-inlined) or as an https:// URL if the model can reach it. Data URLs are the reliable choice inside a lab.

The lab image

The starter code writes a small test image to /tmp/test_img.jpg: a blue square on a yellow background. Your job is to describe it correctly.

Choosing a VLM

We'll use google/gemma-4-31b-it, a multimodal model that handles charts, screenshots and documents and, unlike most vision models, also supports function calling. That second property is what step 4 turns into a lesson. Give it ~800 tokens of headroom.

NVIDIA retired its own hosted Nemotron vision model, and no NVIDIA vision model is currently servable. That is itself the lesson: pin a hosted model in your code and you inherit its retirement date.

Your task

Add to main.py:

  1. Read /tmp/test_img.jpg, base64-encode it, prefix with data:image/jpeg;base64,.
  2. Build a chat request with one text part and one image_url part.
  3. Call google/gemma-4-31b-it with max_tokens=800.
  4. Print the answer.
main.py, the file you edit28 lines
from openai import OpenAI
from PIL import Image, ImageDraw
import base64

# Create a test image: yellow background with a blue square in the middle
def make_image(path="/tmp/test_img.jpg"):
    img = Image.new('RGB', (300, 200), color='yellow')
    ImageDraw.Draw(img).rectangle([100, 60, 200, 140], fill='blue', outline='black', width=3)
    img.save(path, 'JPEG', quality=85)
    return path

def image_to_data_url(path: str) -> str:
    with open(path, 'rb') as f:
        b64 = base64.b64encode(f.read()).decode()
    return f"data:image/jpeg;base64,{b64}"

make_image()
data_url = image_to_data_url("/tmp/test_img.jpg")

client = OpenAI(base_url="http://nim-proxy.labs.svc:8080/v1", api_key="nvapi-inject")
VLM_MODEL = "google/gemma-4-31b-it"

# TODO: describe_image(prompt, data_url) -> str
def describe_image(prompt: str, data_url: str) -> str:
    ______

answer = describe_image("In one sentence: what shape is in the center, what is its color, and what is the background color?", data_url)
print(answer)

Exam domains covered

Multimodal AIAgent DevelopmentNVIDIA Platform Implementation

Frequently asked questions

How do I pass an image to a VLM through the chat completions API?

Swap the content string on the user message for a list of content parts. Each part has a type field: type: "text" for the prompt, type: "image_url" for an image. The image_url can be an https:// URL the model can fetch, or a data URL like data:image/jpeg;base64,<b64> — the base64-inlined form is the reliable choice inside a lab because there's no network fetch to depend on. Nothing else about the request changes — model, messages, temperature, and the response format are the same chat completion you already know.

Can a VLM reason about multiple images in a single turn?

Yes. Put more than one image_url part in the content list of a single user message and the VLM attends to all of them inside one context. Step 2 of this lab puts a 3-circle image next to a 5-circle image and asks which has more; the model sees both and answers from the joint context. This is the primitive behind comparison tasks — A/B screenshots, receipt vs. ledger entry, before/after photos — and it scales up to any number of images the model's context window can hold.

Why does function calling work with one vision model but not the other?

Function calling is a per-model capability, not a blanket feature of multimodal endpoints. google/gemma-4-31b-it ships with tools support, so Step 3 gets back a clean tool_calls with validated arguments. google/gemma-3-12b-it does not: it answers in prose and ignores the schema, so a parser expecting JSON gets free text. Step 4 uses this split deliberately: the agent wrapper you build has to detect which mode is available and fall back to prompt-only JSON when tools aren't supported.

Is prompt-only JSON extraction reliable enough on a VLM?

Less reliable than function calling, but workable for a single well-scoped schema. On the fake receipt in Step 4, google/gemma-3-12b-it asked for JSON in prose produces mostly-correct output most of the time, but you need defensive parsing (strip markdown fences, tolerate trailing commas, null-check every field) and you should treat extraction failures as expected. For complex schemas or production-scale extraction you'd want a VLM that supports tools — or you'd use a VLM to describe the image and a tool-capable text model to do the actual structured extraction.

How is the Step 3 structured extraction different from a text-only tool call?

Mechanically it's identical — same tools parameter, same tool_choice, same function.arguments string on the return — but the model is looking at pixels to fill the schema instead of reading text. That means extraction accuracy depends on image quality and rendering style, not just language. The lab deliberately draws a clean, high-contrast receipt so the vision stage isn't the bottleneck and you can focus on the schema-enforcement pattern. Noisy receipts would add an OCR-style error mode on top.

What's the difference between this VLM lab and the multimodal RAG one?

This lab focuses on VLM core capabilities: captioning, multi-image comparison, and schema-enforced extraction from a single image. The multimodal-rag lab layers a VLM into a retrieval pipeline — the image becomes a query that gets translated into text, retrieved against a product corpus, and grounded in the top-k passages. Here the VLM answers from pixels alone; there the VLM is one component in a larger retrieval-augmented system. The lab order matters: you need the primitives from this one to understand how they plug into the RAG one.

What you'll build in this VLM visual Q&A lab

Vision-Language Models are the fastest-growing surface in production LLM apps — receipt parsing, screenshot triage, document extraction, multimodal agent input — and the tool that separates teams who can ship these features from teams still trying to bolt OCR together. This lab goes from sending a single image into a VLM all the way to a head-to-head comparison between two VLMs on a structured-extraction task, all against NVIDIA NIM endpoints we provision. You finish with working code for single-image Q&A, multi-image reasoning, schema-enforced extraction from a receipt, and a tool-tolerant helper that handles the reality that not every VLM supports function calling.

The technical core is the OpenAI-compatible multimodal content format — text and image_url parts coexist inside a single user message, the VLM fuses them into one context, and the response surface is the exact same chat completion shape text-only calls return. You'll work through base64 data URL payloads (the reliable image-payload choice when you don't want a network fetch in the loop), multi-image reasoning that's just more parts in the list rather than a separate API, and schema-enforced extraction using tools plus tool_choice for a save_receipt function with fields vendor, order_id, line items, subtotal, tax, total. The final step runs the same receipt through the tool-capable model with function calling and through google/gemma-3-12b-it with prompt-only JSON, because the second model ignores the tool schema and answers in prose when you pass tools — a real production split your code needs to handle.

Prerequisites: Python, prior NIM exposure (the react-agent-nim lab works), base64 payload comfort, and basic JSON. No VLM-specific library is assumed — everything goes through the OpenAI Python SDK pointed at our managed NIM proxy, where both google/gemma-4-31b-it and google/gemma-3-12b-it are reached via the same OpenAI-compatible endpoint with no GPU provisioning. About 30 minutes of focused work. You leave with a single-image Q&A call, a multi-image comparison, schema-validated extraction via function calling, and a dual-path extractor that falls back to prompt-only JSON when tools aren't supported — the same defensive shape production VLM code needs.