Part of the AI Infrastructure Engineer Course
Build & submit taskBetaadvanced

Triage a Multi-Tenant GPU Cluster Failing Three Ways at Once

The capstone of the AI Infrastructure Engineer course: a capture from a shared GPU cluster where three teams filed incidents on the same morning. Production inference is being preempted, a training job keeps losing its checkpoints, and an A100 is running at eight percent. The three causes are unrelated and none of them sits on the object that shows the symptom. Diagnose all three from the capture, ship a manifest that fixes each, and write the runbook. Graded against a rubric.

3 hrs

Est. time

5

Outcomes

6

Rubric criteria

65%

Pass score

What you'll learn

Skills you'll have real reps in after shipping this.

The symptom and the cause are rarely the same object
A pod stuck Pending, a job restarting from step zero and an idle GPU are three symptoms. The causes are a PriorityClass field, a StorageClass choice and a device-plugin setting, and none of them appears in the pod that hurts. Triage is the walk from one to the other.
Cluster-wide defaults are a blast radius, not a convenience
globalDefault on a PriorityClass silently rewrites the priority of every workload that did not ask for one, including the customer-facing ones. The same shape recurs with default StorageClasses and default resource limits: a default is a decision applied to everything you forgot about.
Storage tier is a durability decision made at claim time
Node-local scratch is correct for shard-local temporary data and wrong for anything you intend to resume from. The reclaim policy and access mode you pick when you write the PVC decide whether a reschedule costs you a pod or a week of training.
Partitioned GPUs are only capacity if the plugin advertises them
MIG instances that exist on the card but are not advertised by the device plugin are invisible to the scheduler, so the cluster reports full while most of a card idles. Capacity is what the scheduler can see, not what the hardware can do.
Order of operations is part of the fix
Three correct fixes applied in the wrong order still costs you an outage window. Restore the thing with customers behind it first, return capacity before you rely on it, and take the restart-costing fix last.

The scenario

You are on call for ACME Cloud's shared GPU platform. At 09:14 on a Monday you have three tickets open at once and one maintenance window to fix them in.

The serving team says the inference gateway is down to one replica and the rest will not schedule. The research team says their Mistral run has restarted three times and started from step zero each time. The finance dashboard says an eighty-thousand-dollar A100 spent the weekend at eight percent utilisation while the scheduler claimed the cluster was full.

The on-call engineer before you took a capture before the window closed: the JSON that kubectl returns for pods, nodes, events, priority classes and the storage layer, a mig-parted export, and twenty-four hours of DCGM samples. That capture is all you get. The cluster itself is gone.

Three failures, three different subsystems, no shared cause. Fixing one will not fix another, and the object showing the symptom is almost never the object at fault.

Your role

You are the platform engineer who owns this cluster. You work the capture the way you would work a live incident: walk the pod lifecycle from admission to readiness, find the earliest stage where reality diverges from intent, and name the object and field that caused it. Then you ship manifests that a colleague can apply without asking you what you meant, and a runbook that lets the next person on call fix it without you.

Start the task to unlock the full brief

You'll get the step-by-step requirements, setup commands, the 6-criterion grading rubric, tips, and the ability to submit your solution for instant AI grading.

Free to start · submit when you're ready

What this task is

This is the capstone of the AI Infrastructure Engineer course: a build-and-submit incident triage rather than a quiz about Kubernetes. You get the capture from a shared GPU cluster that failed three ways on the same morning, and you produce what a platform team would actually need from you: a machine-readable findings file, one applyable manifest per fix, and a runbook the next person on call can follow.

The three failures are independent and sit in three different subsystems. Production inference is being preempted because a low-priority class was made the cluster-wide default and the serving Deployment never asked for a class of its own. A training job restarts from step zero because its checkpoints live on node-local scratch that cannot follow the pod. An A100 idles at eight percent because seven MIG instances exist on the card and the device plugin advertises one. None of them is visible on the object that shows the symptom.

Everything runs offline on the Python standard library. The starter kit ships the capture in the shapes kubectl returns, an offline kubectl that speaks get and describe over it, a first-pass triage script, and a strict self-check that tells you whether your triage is coherent and grounded before you submit.

Frequently asked questions

Do I need a Kubernetes cluster or a GPU?

No. The cluster in this incident is gone, which is the point: you work the capture the on-call engineer took, exactly as you would after a maintenance window closes. Everything runs offline on the Python standard library, with no cluster, no GPU, no API key and no install.

How is this different from the course labs?

The labs each break one thing in a live environment and check that you fixed it. This breaks three unrelated things at once and does not tell you which subsystems they are in. The work is the diagnosis and the judgement about what to fix first, which is what the labs deliberately isolate you from.

Does the self-check tell me if I am right?

Partly. It confirms your findings are grounded in the capture, that you have covered three distinct subsystems, and that each manifest changes the field your finding blames. It cannot judge your severity calls, your ordering argument or your runbook, and those are a third of the rubric.

Is there one correct set of fixes?

There is one correct set of root causes. There is more than one defensible fix for each: removing the cluster-wide default and pinning the serving workload explicitly are both reasonable for the first failure, and the third accepts either a strategy change or a documented re-partition. The grader rewards a fix that changes the offending field and explains itself.

How long should this take?

About three hours. Most of it is reading the capture rather than writing, and the first pass with the provided triage script takes minutes. If you have finished in under an hour you have probably found two failures rather than three.