An AI answer can accelerate the beginning of research, but it is not evidence by itself. Natural language, precise details, and a confident tone can coexist with a nonexistent source, a source that supports a different conclusion, or information that was true only for an older product version.

NIST's generative AI risk profile treats this kind of false or misleading generation as a confabulation risk. The practical response is not to ask the model to sound more certain. It is to convert the answer into claims that a person can compare with primary evidence.

This guide provides an eight-step process for general research, content production, product comparisons, and workplace decisions. It does not replace qualified review in medicine, law, finance, safety, or other areas where an incorrect claim can cause substantial harm.

1. Break the answer into verifiable claims

Do not label a long answer as entirely true or false. Separate every sentence that presents something as a fact.

For each claim, record:

  • the exact wording produced by the AI
  • the type of claim
  • the primary evidence that would be needed
  • the consequence if it is wrong
  • the relevant date and scope
  • its current verification status

At minimum, distinguish these claim types.

TypeExampleWhat to check first
NumberPrice, percentage, performance, usageUnit, conditions, original data
DateRelease date, policy start dateTime zone, revision, current validity
Product or policyFeature support, retention periodPlan, region, account setting
QuotationA statement attributed to a person or documentOriginal wording and context
CausationA caused BStudy design and alternative explanations
EvaluationBetter, safer, easierCriteria and target user

If one sentence combines several claims, split it again. “This tool is free, stores no data, and is the safest choice for small teams” contains separate claims about price, retention, security, and suitability.

2. Rank claims by risk and freshness

Not every sentence needs the same level of investigation. Verify claims first when an error would cause greater harm or when the information changes frequently.

Give higher priority to claims that:

  • affect a purchase, contract, hiring decision, or public publication
  • concern privacy, security, law, health, or finance
  • involve prices, features, policies, or availability
  • include a specific number, date, or quotation
  • serve as a premise for the rest of the conclusion
  • rely on one vague or unidentified source

A simple ranking method is impact × likelihood of change × difficulty of detection. If two or more factors are high, require explicit human verification before the claim is used.

3. Find the primary source instead of trusting the supplied link

A citation attached to an AI answer is not proof. A model may connect the wrong title to a real URL, present a secondary summary as the original, or invent publication details.

Prefer evidence in this order when it is available:

  1. official product documentation, policies, terms, and release notes
  2. government or standards-body publications
  3. the original paper and underlying data
  4. official corporate filings or announcements
  5. a directly observed product screen or reproducible result
  6. reliable secondary analysis

A DOI can help identify a registered scholarly work, but the existence of a DOI does not prove that a claim is correct. Confirm that the DOI resolves and that its title, authors, publication year, and actual content match what the AI described.

4. Verify the identity and scope of the source

Before relying on the text, establish what the document is and where it applies.

Check:

  • publisher and author
  • original publication and last revision dates
  • document or product version
  • jurisdiction or region
  • applicable plan and account type
  • whether it is official, draft, editorial, or community content
  • whether it contains original evidence or repeats another source

For example, the claim “the service does not use inputs for model training” may differ between consumer accounts, enterprise plans, APIs, regions, and optional settings. A correct link can still support an incorrect conclusion if its scope is ignored.

For time-sensitive material, record a date such as checked 2026-07-19.

5. Test whether the source actually supports the claim

A common verification failure is finding a relevant document that does not say what the answer claims.

Ask:

  • Does the source explicitly state the fact?
  • Did the AI omit a condition, exception, or limitation?
  • Did it turn a possibility into a certainty?
  • Did it turn correlation into causation?
  • Did it expand one feature or test result to every user?
  • Does the surrounding context change the meaning of a quotation?

Store the exact section that supports the conclusion, not just the URL.

Claim: Inputs on the enterprise plan are not used for model training.
Source: Official data policy
Scope: Enterprise plan under specified settings
Result: Supported with conditions
Remaining question: Retention and administrator logging require separate review
Checked: 2026-07-19

6. Cross-check with independent evidence or reproduction

Do not rely on one source for an important claim. However, several articles that repeat the same press release are not independent evidence.

Choose the cross-check method by claim type.

  • Product feature: official documentation plus a test in the actual account
  • Price: official pricing page plus the final purchase screen
  • Statistic: report narrative plus the original data table
  • Research result: methods, results, and limitations rather than only the abstract
  • Calculation: repeat it from the original inputs and formula
  • Quotation: original video, transcript, meeting record, or document
  • Policy: current document plus revision history

When sources disagree, do not average them or choose the majority automatically. Compare dates, definitions, samples, regions, and test conditions first.

7. Preserve uncertainty as a status

Verification will not remove every uncertainty. Filling the remaining gaps with a plausible guess is more dangerous than marking them clearly.

A claim ledger can use these statuses.

StatusMeaningUse
VerifiedA primary source directly supports itRecord conditions and checked date
Conditionally verifiedTrue only for a plan, region, or scenarioState the scope in the final text
ConflictingCredible sources disagreeExplain the likely reason
UnverifiedNo adequate primary evidence foundDo not publish as fact
VolatilePrice, policy, or feature changes quicklyAdd a date and recheck notice
OpinionA judgment rather than a measured factPublish the evaluation criteria

Important uncertainty should remain visible to the reader. “We could not verify this” or “the current documentation does not clarify the condition” defines the evidence boundary rather than weakening the work.

8. Keep human approval and an audit trail as separate steps

Asking the same AI to critique its own answer is not independent verification. A second model may expose additional problems, but it still does not replace checking the primary evidence.

The final reviewer should confirm:

  • every important statement appears in the claim ledger
  • the source text and the published claim match
  • dates, regions, plans, settings, and exceptions are visible
  • no unresolved claim remains essential to the conclusion
  • quotations and numbers preserve their original meaning
  • there is a correction path after publication or use

The NIST AI Risk Management Framework describes risk management as continuous work across Govern, Map, Measure, and Manage. The GAO accountability framework similarly combines governance, data, performance, and monitoring. In practice, a record of who checked what, when, and under which version is more useful than a vague statement that the answer was “fact-checked.”

Practical verification matrix

FieldQuestionEvidence to keep
ClaimWhat exactly is presented as true?Exact claim text
RiskWhat happens if it is wrong?Impact rating
Primary sourceWhat is the closest official evidence?URL and document title
IdentityAre author, publisher, and version correct?Publication and revision dates
ScopeWhich plan, region, and conditions apply?Conditions and exceptions
EntailmentDoes the evidence support this conclusion?Relevant section
Cross-checkWas it reproduced or independently confirmed?Second evidence item
StatusVerified, conditional, conflicting, or unverified?Status and reason
ApprovalWho authorized final use?Reviewer and date

Common mistakes

  • Copying an AI-generated bibliography without opening each source
  • Reading a search-result snippet instead of the original document
  • Counting multiple retellings of one source as independent confirmation
  • Ignoring the document date or product version
  • Treating a DOI, institutional logo, or professional tone as proof
  • Discarding contradictory evidence because it complicates the conclusion
  • Hiding an unverified statement behind words such as “generally” or “typically”
  • Confusing the model's confidence with the probability that a claim is true

Short pre-publication checklist

  • Did you separate the answer into individual claims?
  • Did you verify the high-risk and fast-changing claims first?
  • Did you open the primary source and read the relevant section?
  • Does the source directly support the published conclusion?
  • Did you record the date, region, plan, and exceptions?
  • Did you independently cross-check important claims?
  • Did you mark unverified, conflicting, and volatile information?
  • Did a person approve the final use or publication?

When you are evaluating the tool itself, use this process with 7 Criteria to Check Before Choosing an AI Tool. Choosing a capable tool and verifying an individual answer are related but separate decisions. A useful product can still produce a claim that requires fresh evidence.

Sources reviewed

Sources checked: 2026-07-19