Introduction
This AI draft checklist gives reviewers a more concrete test. It scores seven parts of format readiness: task fit, source fidelity, format completeness, information hierarchy, decision clarity, edit burden, and handoff independence.
The checklist is meant for briefs, updates, summaries, FAQs, and other work that has a clear reader and job. It does not measure model performance. It helps a person decide what kind of draft they have and what must happen before it moves forward.
How the AI draft checklist works
Score each check from 0 to 2.
| Score | Meaning | Review consequence |
|---|---|---|
| 0 | Missing or wrong in a way that blocks use | Stop and repair the structure or evidence |
| 1 | Present, but still needs structural revision | Keep it as a working draft |
| 2 | Usable with local edits and normal subject review | Move to the next review gate |
Add the seven scores for a total out of 14.
| Total | Draft state | What to do next |
|---|---|---|
| 0 to 5 | Raw output | Reframe the task or rebuild the format |
| 6 to 10 | Working draft | Resolve the failed checks before handoff |
| 11 to 14 | Handoff candidate | Continue with factual, editorial, legal, or specialist review as the work requires |
The total is useful, but it is not the only rule. Task fit and source fidelity are hard gates. A draft with a 0 on either one is not a handoff candidate, even if the other scores are high.
This matters because generated prose can be confidently wrong or can diverge from its input. The NIST Generative AI Profile calls this risk confabulation. A good layout cannot compensate for an unsupported claim, and a faithful summary cannot compensate for answering the wrong task.
The scorecard is an editorial framework developed by the FormaLM Team and first published on August 3, 2026. The score bands are working thresholds, not results from a scientific benchmark. Use the downloadable CSV scorecard if you want to adapt the checks to a repeated review process.
The seven format-readiness tests
1. Task fit
Ask: does the draft do the job it was asked to do?
A response can be well written and still fail because it chose the wrong reader, level of detail, or outcome. A project brief is not a meeting summary. An executive update is not a transcript with cleaner sentences. A customer FAQ should answer likely questions, not preserve the order of the source notes.
Give a 0 when the draft needs a new frame. Give a 1 when the basic task is visible but the audience or purpose still drifts. Give a 2 when the intended reader and use are clear throughout.
2. Source fidelity
Ask: can each material claim be traced to the supplied source or a cited source?
Source fidelity is more than factual tone. Check names, dates, quantities, conclusions, and attributed opinions against the material that supports them. Mark uncertain or missing information instead of asking polished prose to hide the gap.
Give a 0 when the draft invents, contradicts, or cannot support material claims. Give a 1 when it is mostly grounded but still has gaps to resolve. Give a 2 when claims are traceable and unresolved points are explicit.
If your input is still a loose set of links or notes, a source-to-structure summarizer can help separate source material from the final format. It does not remove the need to check the claims.
3. Format completeness
Ask: are the parts required by the chosen format present?
Every useful format creates an expectation. A brief may need an objective, audience, scope, constraints, and open questions. A weekly update may need progress, decisions, blockers, and next steps. Missing one of those parts can make the document feel oddly thin even when the sentences are strong.
Give a 0 when missing sections block use. Give a 1 when the core is present but one or more sections still need structural work. Give a 2 when the required parts are present and proportionate.
The project brief template shows how format requirements change the review. You are checking whether the draft behaves like a brief, not whether it contains enough words.
4. Information hierarchy
Ask: can the reader tell what matters first and why?
Generated drafts often flatten the source. Every point gets a similar paragraph, minor detail receives the same emphasis as the decision, and repeated ideas survive because each sentence sounds plausible on its own.
Give a 0 when the order is confusing or repetitive. Give a 1 when the main point exists but the emphasis still needs work. Give a 2 when the reader can scan the priority, context, and supporting detail without reconstructing the logic.
5. Decision clarity
Ask: does the draft make the requested decision or next action clear?
Not every document needs an action item. When the format does, the action should be operational. "Improve onboarding" is a direction. "Mia will revise the first-run email before Friday's review" is an action with an owner and boundary.
Give a 0 when the draft ends without a usable decision or action. Give a 1 when the direction is implied but still needs interpretation. Give a 2 when the decision, owner, or next action is explicit where the format requires it.
6. Edit burden
Ask: can the remaining review stay local, or does the document need a structural rewrite?
Local edits fix a sentence, label, example, or small omission. Structural edits change the audience, order, section logic, evidence base, or document type. That distinction is more useful than asking whether the draft "needs editing," because nearly every draft does.
Give a 0 when a reviewer must rebuild the document. Give a 1 when several sections need rewriting or reordering. Give a 2 when only local edits and normal subject review remain.
7. Handoff independence
Ask: can the intended reader use the draft without the author in the room?
Private context often survives inside generated work. A phrase may make sense only to the person who supplied the notes. A pronoun may refer to an unstated project. A recommendation may depend on a constraint that never reached the page.
Give a 0 when the document depends on missing context. Give a 1 when it becomes usable after a short explanation. Give a 2 when it stands on its own for the intended handoff.
This is also the final test in the notes-to-brief workflow. A document becomes shareable when its structure carries the context that used to live only in the author's head.
A worked example: product update brief
Imagine a product manager asks a generative tool to turn meeting notes into a one-page update for leadership. The output has a clear summary, three progress bullets, a polished explanation of a delay, and a next-steps section.
During review, the product manager notices two problems. The draft says the release is "on track," but the notes only say that engineering expects to confirm timing next week. It also leaves the launch decision implicit.
The score might look like this:
| Check | Score | Review note |
|---|---|---|
| Task fit | 2 | It is a compact leadership update |
| Source fidelity | 0 | "On track" is not supported by the notes |
| Format completeness | 2 | Summary, progress, risk, and next step are present |
| Information hierarchy | 2 | The delay and current progress are easy to scan |
| Decision clarity | 1 | The approval needed from leadership is implied |
| Edit burden | 1 | Two sections need more than copyediting |
| Handoff independence | 2 | Leadership can understand the rest without more context |
| Total | 10 | Working draft; source fidelity blocks handoff |
The total alone says "working draft." The hard gate gives the sharper answer: do not hand it off until the unsupported status claim is corrected. After that, make the approval request explicit and score the draft again.
This is why the checklist uses separate dimensions. One smooth paragraph should not cancel out one serious evidence failure.
What the score does not approve
A 14 out of 14 means the draft is ready for the next human gate. It does not mean the content is true in every respect, legally approved, safe for a regulated use, on brand, accessible, or ready to publish without review.
Add domain-specific gates when the work carries more risk. A medical summary may need clinical review. A financial claim may need compliance review. A public product announcement may need legal, security, accessibility, and brand checks. The scorecard does not replace those responsibilities.
It also does not compare models. If you want to use it for evaluation, define a stable task set, record the source material, preserve reviewer notes, and check agreement between reviewers. Without that method, a single score is a workflow note, not performance evidence.
Use the checklist as a workflow gate
Run the checklist when the format is visible but before the draft enters a costly review cycle.
For a one-off document, one reviewer can score it in a few minutes and add notes only to failed checks. For repeated work, keep the seven fields beside the draft and track which failures recur. If format completeness is often low, the workflow may need a better template. If source fidelity fails, the source boundary needs to become clearer before generation. If handoff independence stays weak, the input may contain too much private context.
The checklist works best as a stop condition, not a decorative score. A draft moves forward when the hard gates pass and the remaining edits match the risk of the task.
That is the practical difference between generated text and usable work. The finish line is not fluency. It is a draft whose evidence, structure, and handoff all hold together.
