An AI answer can accelerate the beginning of research, but it is not evidence by itself. Natural language, precise details, and a confident tone can coexist with a nonexistent source, a source that supports a different conclusion, or information that was true only for an older product version.
NIST's generative AI risk profile treats this kind of false or misleading generation as a confabulation risk. The practical response is not to ask the model to sound more certain. It is to convert the answer into claims that a person can compare with primary evidence.
This guide provides an eight-step process for general research, content production, product comparisons, and workplace decisions. It does not replace qualified review in medicine, law, finance, safety, or other areas where an incorrect claim can cause substantial harm.
1. Break the answer into verifiable claims
Do not label a long answer as entirely true or false. Separate every sentence that presents something as a fact.
For each claim, record:
- the exact wording produced by the AI
- the type of claim
- the primary evidence that would be needed
- the consequence if it is wrong
- the relevant date and scope
- its current verification status
At minimum, distinguish these claim types.
| Type | Example | What to check first |
|---|---|---|
| Number | Price, percentage, performance, usage | Unit, conditions, original data |
| Date | Release date, policy start date | Time zone, revision, current validity |
| Product or policy | Feature support, retention period | Plan, region, account setting |
| Quotation | A statement attributed to a person or document | Original wording and context |
| Causation | A caused B | Study design and alternative explanations |
| Evaluation | Better, safer, easier | Criteria and target user |
If one sentence combines several claims, split it again. “This tool is free, stores no data, and is the safest choice for small teams” contains separate claims about price, retention, security, and suitability.
2. Rank claims by risk and freshness
Not every sentence needs the same level of investigation. Verify claims first when an error would cause greater harm or when the information changes frequently.
Give higher priority to claims that:
- affect a purchase, contract, hiring decision, or public publication
- concern privacy, security, law, health, or finance
- involve prices, features, policies, or availability
- include a specific number, date, or quotation
- serve as a premise for the rest of the conclusion
- rely on one vague or unidentified source
A simple ranking method is impact × likelihood of change × difficulty of detection. If two or more factors are high, require explicit human verification before the claim is used.
3. Find the primary source instead of trusting the supplied link
A citation attached to an AI answer is not proof. A model may connect the wrong title to a real URL, present a secondary summary as the original, or invent publication details.
Prefer evidence in this order when it is available:
- official product documentation, policies, terms, and release notes
- government or standards-body publications
- the original paper and underlying data
- official corporate filings or announcements
- a directly observed product screen or reproducible result
- reliable secondary analysis
A DOI can help identify a registered scholarly work, but the existence of a DOI does not prove that a claim is correct. Confirm that the DOI resolves and that its title, authors, publication year, and actual content match what the AI described.
4. Verify the identity and scope of the source
Before relying on the text, establish what the document is and where it applies.
Check:
- publisher and author
- original publication and last revision dates
- document or product version
- jurisdiction or region
- applicable plan and account type
- whether it is official, draft, editorial, or community content
- whether it contains original evidence or repeats another source
For example, the claim “the service does not use inputs for model training” may differ between consumer accounts, enterprise plans, APIs, regions, and optional settings. A correct link can still support an incorrect conclusion if its scope is ignored.
For time-sensitive material, record a date such as checked 2026-07-19.
5. Test whether the source actually supports the claim
A common verification failure is finding a relevant document that does not say what the answer claims.
Ask:
- Does the source explicitly state the fact?
- Did the AI omit a condition, exception, or limitation?
- Did it turn a possibility into a certainty?
- Did it turn correlation into causation?
- Did it expand one feature or test result to every user?
- Does the surrounding context change the meaning of a quotation?
Store the exact section that supports the conclusion, not just the URL.
Claim: Inputs on the enterprise plan are not used for model training.
Source: Official data policy
Scope: Enterprise plan under specified settings
Result: Supported with conditions
Remaining question: Retention and administrator logging require separate review
Checked: 2026-07-19
6. Cross-check with independent evidence or reproduction
Do not rely on one source for an important claim. However, several articles that repeat the same press release are not independent evidence.
Choose the cross-check method by claim type.
- Product feature: official documentation plus a test in the actual account
- Price: official pricing page plus the final purchase screen
- Statistic: report narrative plus the original data table
- Research result: methods, results, and limitations rather than only the abstract
- Calculation: repeat it from the original inputs and formula
- Quotation: original video, transcript, meeting record, or document
- Policy: current document plus revision history
When sources disagree, do not average them or choose the majority automatically. Compare dates, definitions, samples, regions, and test conditions first.
7. Preserve uncertainty as a status
Verification will not remove every uncertainty. Filling the remaining gaps with a plausible guess is more dangerous than marking them clearly.
A claim ledger can use these statuses.
| Status | Meaning | Use |
|---|---|---|
| Verified | A primary source directly supports it | Record conditions and checked date |
| Conditionally verified | True only for a plan, region, or scenario | State the scope in the final text |
| Conflicting | Credible sources disagree | Explain the likely reason |
| Unverified | No adequate primary evidence found | Do not publish as fact |
| Volatile | Price, policy, or feature changes quickly | Add a date and recheck notice |
| Opinion | A judgment rather than a measured fact | Publish the evaluation criteria |
Important uncertainty should remain visible to the reader. “We could not verify this” or “the current documentation does not clarify the condition” defines the evidence boundary rather than weakening the work.
8. Keep human approval and an audit trail as separate steps
Asking the same AI to critique its own answer is not independent verification. A second model may expose additional problems, but it still does not replace checking the primary evidence.
The final reviewer should confirm:
- every important statement appears in the claim ledger
- the source text and the published claim match
- dates, regions, plans, settings, and exceptions are visible
- no unresolved claim remains essential to the conclusion
- quotations and numbers preserve their original meaning
- there is a correction path after publication or use
The NIST AI Risk Management Framework describes risk management as continuous work across Govern, Map, Measure, and Manage. The GAO accountability framework similarly combines governance, data, performance, and monitoring. In practice, a record of who checked what, when, and under which version is more useful than a vague statement that the answer was “fact-checked.”
Practical verification matrix
| Field | Question | Evidence to keep |
|---|---|---|
| Claim | What exactly is presented as true? | Exact claim text |
| Risk | What happens if it is wrong? | Impact rating |
| Primary source | What is the closest official evidence? | URL and document title |
| Identity | Are author, publisher, and version correct? | Publication and revision dates |
| Scope | Which plan, region, and conditions apply? | Conditions and exceptions |
| Entailment | Does the evidence support this conclusion? | Relevant section |
| Cross-check | Was it reproduced or independently confirmed? | Second evidence item |
| Status | Verified, conditional, conflicting, or unverified? | Status and reason |
| Approval | Who authorized final use? | Reviewer and date |
Common mistakes
- Copying an AI-generated bibliography without opening each source
- Reading a search-result snippet instead of the original document
- Counting multiple retellings of one source as independent confirmation
- Ignoring the document date or product version
- Treating a DOI, institutional logo, or professional tone as proof
- Discarding contradictory evidence because it complicates the conclusion
- Hiding an unverified statement behind words such as “generally” or “typically”
- Confusing the model's confidence with the probability that a claim is true
Short pre-publication checklist
- Did you separate the answer into individual claims?
- Did you verify the high-risk and fast-changing claims first?
- Did you open the primary source and read the relevant section?
- Does the source directly support the published conclusion?
- Did you record the date, region, plan, and exceptions?
- Did you independently cross-check important claims?
- Did you mark unverified, conflicting, and volatile information?
- Did a person approve the final use or publication?
When you are evaluating the tool itself, use this process with 7 Criteria to Check Before Choosing an AI Tool. Choosing a capable tool and verifying an individual answer are related but separate decisions. A useful product can still produce a claim that requires fresh evidence.
Sources reviewed
- NIST Generative AI Profile (opens in a new window)
- NIST AI Risk Management Framework 1.0 (opens in a new window)
- U.S. GAO AI Accountability Framework (opens in a new window)
- Crossref guidance for verifying DOI registration (opens in a new window)
Sources checked: 2026-07-19