Step 1: Send one image and ask a question
Calling a VLM
Vision-Language Models use the same OpenAI-compatible chat endpoint as text-only models, but the content field inside a user message becomes a list of content parts:
{
"role": "user",
"content": [
{"type": "text", "text": "What's in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,<...>"}},
]
}
You can pass an image as a data URL (base64-inlined) or as an https:// URL if the model can reach it. Data URLs are the reliable choice inside a lab.
The lab image
The starter code writes a small test image to /tmp/test_img.jpg: a blue square on a yellow background. Your job is to describe it correctly.
Choosing a VLM
We'll use google/gemma-4-31b-it, a multimodal model that handles charts, screenshots and documents and, unlike most vision models, also supports function calling. That second property is what step 4 turns into a lesson. Give it ~800 tokens of headroom.
NVIDIA retired its own hosted Nemotron vision model, and no NVIDIA vision model is currently servable. That is itself the lesson: pin a hosted model in your code and you inherit its retirement date.
Your task
Add to main.py:
- Read
/tmp/test_img.jpg, base64-encode it, prefix withdata:image/jpeg;base64,. - Build a chat request with one
textpart and oneimage_urlpart. - Call
google/gemma-4-31b-itwithmax_tokens=800. - Print the answer.
main.py, the file you edit28 lines
from openai import OpenAI
from PIL import Image, ImageDraw
import base64
# Create a test image: yellow background with a blue square in the middle
def make_image(path="/tmp/test_img.jpg"):
img = Image.new('RGB', (300, 200), color='yellow')
ImageDraw.Draw(img).rectangle([100, 60, 200, 140], fill='blue', outline='black', width=3)
img.save(path, 'JPEG', quality=85)
return path
def image_to_data_url(path: str) -> str:
with open(path, 'rb') as f:
b64 = base64.b64encode(f.read()).decode()
return f"data:image/jpeg;base64,{b64}"
make_image()
data_url = image_to_data_url("/tmp/test_img.jpg")
client = OpenAI(base_url="http://nim-proxy.labs.svc:8080/v1", api_key="nvapi-inject")
VLM_MODEL = "google/gemma-4-31b-it"
# TODO: describe_image(prompt, data_url) -> str
def describe_image(prompt: str, data_url: str) -> str:
______
answer = describe_image("In one sentence: what shape is in the center, what is its color, and what is the background color?", data_url)
print(answer)