Build & submit taskBetaintermediate

Fact-Check a Claude Research Brief and Ship a Claim Register

Ask Claude for a research brief, then audit it claim by claim: extract every checkable factual assertion into a register, verify each one against a primary source, mark it verified, unsupported, or contradicted, and set the threshold at which a human has to review before the work leaves the building. Includes adapting the corrected brief for a second audience. No coding required.

2.5 hrs

Est. time

5

Outcomes

8

Rubric criteria

65%

Pass score

What you'll learn

Skills you'll have real reps in after shipping this.

Fluency is not evidence
Confident phrasing and clean structure carry no information about whether a claim is true. The audit has to be mechanical because your intuition is calibrated on writing quality rather than accuracy.
Extract before you verify
Pulling assertions out of the prose into a register is what makes verification tractable. Reading for errors in flowing text is how the two wrong figures got through in the first place.
Unsupported is its own verdict
A claim you could not confirm is different from one you disproved, and collapsing the two either overstates the error rate or quietly ships something unverified.
Thresholds beat vigilance
Deciding in advance which claim categories always need human sign-off scales across a team. Relying on everyone being careful does not.
Adaptation is a validation step
Rewriting for a second audience surfaces claims that were only ever plausible because the original phrasing was vague.

See how it works

Claims traced back to sources

grounding and citations
click a [n] in the answer
context (retrieved chunks)
[1]
The v3 billing API allows 120 requests per minute per API key.
billing-v3.md
[2]
Exceeding a rate limit returns HTTP 429 with a Retry-After header.
errors.md
[3]
API keys are scoped per project and can be rotated in the dashboard.
auth.md
answer
The v3 billing API permits 120 requests per minute per key, and going over returns a 429 with a Retry-After header. Limits reset at midnight UTC [no source].
"Limits reset at midnight UTC" appears in no chunk. It is an ungrounded claim the model invented.
Every claim should trace to a chunk. Retrieval found the evidence; grounding is the prompt discipline that makes the model use it and only it. You number the chunks, instruct the model to answer from the context and cite the number behind each claim, and now every sentence is checkable: click a citation and you land on its source. The sentence with no citation is exactly the failure mode to hunt, a confident assertion the context never supported.

The audit turns flowing prose into a register where every assertion points at a source or is marked as pointing at nothing. That mapping is what makes an error rate measurable instead of a matter of impression.

The scenario

A colleague forwards a Claude-written market brief to a client without reading it closely. Two of the figures in it are wrong, one cited report does not exist under that title, and the client notices before you do. The output was fluent, well organized, and confidently worded, which is exactly why nobody stopped to check it. Fluency is not evidence, and a hallucination is a claim the model generated because it fit the pattern of the surrounding text rather than because it came from a source.

The habit that prevents this is mechanical rather than intuitive. You pull every checkable assertion out of the prose into a register, verify each one against a primary source you can link, and record the verdict next to the claim. Claims that cannot be verified do not get quietly deleted, they get marked and escalated. This task makes you run that audit end to end on a brief you generate yourself, then decide in advance which categories of claim always require a human before the work ships.

Your role

You are the person who signs off before AI-assisted research reaches a client or an executive. Your deliverable is a claim register for a Claude-generated brief, with a verdict and a source for every factual assertion, plus a written human-review threshold your team could adopt.

Start the task to unlock the full brief

You'll get the step-by-step requirements, setup commands, the 8-criterion grading rubric, tips, and the ability to submit your solution for instant AI grading.

Free to start · submit when you're ready

What you'll build in this output validation task

This is a build-and-submit task rather than a guided lab, and it requires no coding. You generate a research brief with Claude, then audit it the way a careful analyst audits any draft that is about to reach a client: every checkable assertion goes into a register, gets a verdict of verified, unsupported, or contradicted, and carries a primary source link or a note explaining why none exists.

The point is to build a habit that does not depend on being alert. You tally the verdicts into an error rate, identify what the failing claims have in common, and then write a human-review threshold naming the claim categories that always require sign-off before work leaves your team. The final step adapts the corrected brief for a second audience, which tends to expose claims that only read as solid because the original wording was vague.

Grading is rubric-based and explainable. Your document is scored against weighted criteria covering the register, the sourcing, the error-rate tally, the pattern analysis, and the review threshold, with per-criterion feedback quoted from your submission. The pass threshold is 65 percent and you can resubmit. Output evaluation and validation is the heaviest scored domain on the Claude Certified Associate Foundations exam.

Frequently asked questions

Do I need coding or API access?

No. The whole task runs in claude.ai, including the free tier, and the deliverable is a Markdown or PDF document. The Claude Certified Associate Foundations certification is built for professionals with no software development or API experience.

What makes a claim checkable?

A third party could confirm or refute it without knowing your opinion. Figures, dates, named organizations, and cited report titles qualify. Statements like "the market is evolving quickly" do not, and a register full of those cannot produce a meaningful error rate.

What if the brief turns out to be entirely accurate?

Report that, with the sources that confirmed it. A clean result is a real finding when every verdict is backed by a link. The rubric rewards the auditability of your verdicts rather than the number of errors you found.

Why separate unsupported from contradicted?

They call for different responses. A contradicted claim is wrong and must be corrected. An unsupported claim might be true but has no source behind it, so it either gets sourced or it gets escalated. Collapsing the two either inflates your error rate or ships something unverified.

What counts as a complete submission?

One Markdown, text, or PDF file with the exact prompt and verbatim brief, a register of twelve or more checkable claims each carrying a verdict and a source or failed-search note, a verdict tally with an error rate, a pattern analysis of the failures, a justified human-review threshold, and the corrected brief adapted for a second audience.