Build & submit taskBetabeginner

Route Eight Real Tasks to the Right Claude Feature and Model Tier

Build a routing matrix that sends eight recurring work tasks to a specific Claude product feature and model tier, with the cost, speed, and quality tradeoff written out for each. Then test three of your routings head to head on the same input and revise the matrix where the evidence disagrees with you. Includes a context-management plan for long-running work. No coding required.

2 hrs

Est. time

4

Outcomes

8

Rubric criteria

65%

Pass score

What you'll learn

Skills you'll have real reps in after shipping this.

Surfaces have jobs
Projects persist instructions and knowledge across conversations, Artifacts give you an iterable work product beside the chat, and research mode gathers before answering. Matching the surface to the job removes work that repeated setup was silently costing you.
The heaviest tier is not the default answer
Reaching for maximum capability on every task spends cost and latency on work that a lighter tier completes just as well. Knowing where the lighter tier clears the bar is the actual skill.
Test the calls you are unsure about
A routing matrix built purely from reasoning encodes your assumptions. Running the same input two ways turns a few of those assumptions into findings.
Context is a resource you manage
Long conversations accumulate material that stops helping. Deciding in advance when to restart, when to summarize, and what to persist keeps quality steady across work that outlives one thread.

See how it works

Routing work to the right destination

supervisor routing
→ research
incoming request
research
3
matches
finance
1
match
comms
0
matches
Send it to the worker that fits best. A supervisor reads the request and hands it to the specialist built for it; the workers never talk to each other, everything routes through the center. Here that decision is keyword overlap: the worker matching the most of the request wins, not the first to match at all, so a query touching several areas goes to its strongest fit. And a request that matches no specialist returns no_match, so the supervisor can fall back rather than route nonsense to whoever happens to be first.

Every task carries a different mix of volume, difficulty, and lifespan. The matrix makes that routing explicit so the choice stops depending on whichever surface happened to be open.

The scenario

Most teams settle on one Claude habit and apply it to everything: a single chat thread, whatever model is selected by default, and no thought about which surface fits the job. That works until it does not. A recurring weekly report gets rebuilt from scratch every time because nobody set up a Project to hold the context. A quick reformatting job runs on the slowest, most expensive tier for no gain. A long research thread degrades because the conversation outgrew what could usefully be held in context and nobody restarted it.

Choosing well requires knowing what each surface is for. Projects persist instructions and knowledge across conversations. Artifacts give you a document or piece of work you can iterate on beside the chat rather than inside it. Research mode goes and gathers material before answering. Model tiers trade cost and speed against depth, with the lighter tiers suited to high-volume, well-specified work and the heavier ones earning their cost on genuinely hard reasoning. This task makes you commit to a routing for eight real tasks and then check three of those commitments against an actual head-to-head run, because a routing matrix nobody tested is a set of assumptions in a table.

Your role

You are the person your team asks which Claude setup to use for a given job. Your deliverable is a routing matrix for eight recurring tasks, three head-to-head comparisons that test your own calls, and a context-management plan for work that outlives a single conversation.

Start the task to unlock the full brief

You'll get the step-by-step requirements, setup commands, the 8-criterion grading rubric, tips, and the ability to submit your solution for instant AI grading.

Free to start · submit when you're ready

What you'll build in this Claude model selection task

This is a build-and-submit task rather than a guided lab, and it requires no coding. You take eight recurring tasks from your own work and route each one to a specific Claude product feature and model tier, writing out the cost, speed, and quality tradeoff behind every assignment instead of recording the choice on its own.

The part that makes it more than a table is the testing. You pick three routings you were genuinely unsure about, run the same input under both options, paste the verbatim output from each side, and revise the matrix wherever the evidence disagrees with your original call. At least one test has to conclude that the lighter or cheaper option was good enough, with the bar for good enough set before you looked. A context-management plan covering when to restart, when to summarize, and what to persist into a Project closes the document.

Grading is rubric-based and explainable. Your submission is scored against weighted criteria covering coverage, tradeoff reasoning, the head-to-head evidence, the evidence-driven revision, and the context plan, with per-criterion feedback quoted from your document. The pass threshold is 65 percent and you can resubmit. Product and model selection is a scored domain on the Claude Certified Associate Foundations exam.

Frequently asked questions

Do I need an API key or any coding?

No. Everything runs in claude.ai. You do need a plan that lets you switch between model tiers and create a Project, since the head-to-head tests compare tiers directly. The deliverable is a Markdown or PDF document.

How do I choose which routings to test head to head?

Pick the three you were genuinely torn about. Testing a call you were already confident in confirms what you knew and produces no revision, and the rubric rewards the assumptions your evidence actually overturned.

Why require a case where the lighter model was good enough?

Defaulting to the most capable tier for everything spends cost and latency on work that finishes just as well on a lighter one. Knowing where the bar sits is the skill the domain tests, and it only becomes real when you have set the bar in advance and checked against it.

Model tiers and pricing keep changing. How should I handle that?

Check Anthropic's own models and pricing pages while you work and note the date you checked. The rubric asks for a dated check rather than a specific answer, because the routing logic is what transfers.

What counts as a complete submission?

One Markdown, text, or PDF file with an eight-task routing matrix assigning a feature and tier to each with tradeoff reasoning, every product feature used at least once, three head-to-head tests with verbatim output from both sides and a verdict, at least one lighter-option-sufficient finding, recorded revisions where tests changed your mind, and a context-management plan.