Triage a Multi-Tenant GPU Cluster Failing Three Ways at Once
The capstone of the AI Infrastructure Engineer course: a capture from a shared GPU cluster where three teams filed incidents on the same morning. Production inference is being preempted, a training job keeps losing its checkpoints, and an A100 is running at eight percent. The three causes are unrelated and none of them sits on the object that shows the symptom. Diagnose all three from the capture, ship a manifest that fixes each, and write the runbook. Graded against a rubric.
3 hrs
Est. time
5
Outcomes
6
Rubric criteria
65%
Pass score
What you'll learn
Skills you'll have real reps in after shipping this.
The scenario
You are on call for ACME Cloud's shared GPU platform. At 09:14 on a Monday you have three tickets open at once and one maintenance window to fix them in.
The serving team says the inference gateway is down to one replica and the rest will not schedule. The research team says their Mistral run has restarted three times and started from step zero each time. The finance dashboard says an eighty-thousand-dollar A100 spent the weekend at eight percent utilisation while the scheduler claimed the cluster was full.
The on-call engineer before you took a capture before the window closed: the JSON that kubectl returns for pods, nodes, events, priority classes and the storage layer, a mig-parted export, and twenty-four hours of DCGM samples. That capture is all you get. The cluster itself is gone.
Three failures, three different subsystems, no shared cause. Fixing one will not fix another, and the object showing the symptom is almost never the object at fault.
Your role
You are the platform engineer who owns this cluster. You work the capture the way you would work a live incident: walk the pod lifecycle from admission to readiness, find the earliest stage where reality diverges from intent, and name the object and field that caused it. Then you ship manifests that a colleague can apply without asking you what you meant, and a runbook that lets the next person on call fix it without you.
Start the task to unlock the full brief
You'll get the step-by-step requirements, setup commands, the 6-criterion grading rubric, tips, and the ability to submit your solution for instant AI grading.
Free to start · submit when you're ready
Learning resources
What this task is
This is the capstone of the AI Infrastructure Engineer course: a build-and-submit incident triage rather than a quiz about Kubernetes. You get the capture from a shared GPU cluster that failed three ways on the same morning, and you produce what a platform team would actually need from you: a machine-readable findings file, one applyable manifest per fix, and a runbook the next person on call can follow.
The three failures are independent and sit in three different subsystems. Production inference is being preempted because a low-priority class was made the cluster-wide default and the serving Deployment never asked for a class of its own. A training job restarts from step zero because its checkpoints live on node-local scratch that cannot follow the pod. An A100 idles at eight percent because seven MIG instances exist on the card and the device plugin advertises one. None of them is visible on the object that shows the symptom.
Everything runs offline on the Python standard library. The starter kit ships the capture in the shapes kubectl returns, an offline kubectl that speaks get and describe over it, a first-pass triage script, and a strict self-check that tells you whether your triage is coherent and grounded before you submit.