How Accurate Are AI Citations? ChatGPT, Gemini and AI Search Explained

Research reviewed

How accurate are AI citations?

AI citations are useful starting points, but a citation is not proof that an answer is correct. Accuracy depends on whether the source exists, whether the link identifies the right source, whether the source supports the claim, and whether the answer omits important evidence. No single published percentage describes all of those checks across today's AI products.

A link can work and still be misleading. It may lead to a real page that discusses the topic without supporting the number, date or conclusion attached to it.

The safest habit is to verify the claim-source relationship before reusing an answer in an article, business decision or other consequential work.

Five ways an AI citation can fail

FailureExampleWhat to check
Fabricated referenceA convincing paper title with no matching publicationPublisher record, DOI or authoritative catalogue
Broken or wrong URLA link opens an error page or unrelated articleExact destination and redirect chain
Incorrect attributionA copied article is credited as the original reportingOriginal publisher and publication history
Claim-source mismatchA study reports correlation but the answer claims causationRelevant passage, method and limitations
Incomplete evidenceEvery listed reference is real, but key evidence is missingSearch coverage and alternative sources

These failures require different remedies. Fixing a broken URL will not repair a claim that the underlying study never supported.

What public studies actually measured

We compared three research designs rather than combining their percentages into a misleading engine league table.

StudyTask and unitReported resultWhat it cannot establish
Liu, Zhang and Liang, 2023Human evaluation of generated sentences and their citations across four then-current engines51.5% of sentences fully supported; 74.5% of citations supported the associated sentenceAccuracy of today's ChatGPT or Gemini models
Tow Center, March 2025Identify news article, publisher and URL from excerpts; 1,600 queries across eight toolsMore than 60% of queries received incorrect answersError rate for every open-ended AI answer
Liu and colleagues, August 2026 preprintRetrieve studies for 20 medical-review questions; 720 responsesMean included-study recall of 39.2%Citation fabrication rate or general marketing-answer accuracy

Sources and study dates matter.[1][2][3] The first two are historical audits. The third is a recent, domain-specific preprint, not a universal consumer-product ranking.

Our analysis: the denominator changes the meaning

A 74.5% citation-support rate and a 51.5% sentence-support rate can coexist. The former evaluates supplied citations; the latter asks whether generated sentences have full support. An answer can include some good references and still contain unsupported claims.

Similarly, 39.2% recall in a study-retrieval task concerns how much of an expert-selected study set was found. It does not mean that 60.8% of the citations were fabricated.

This comparability audit is CiteLadder's contribution. The underlying observations belong to the cited researchers. We have not conducted a new 10,000-citation test, and we do not present these different studies as a trend line showing that one product improved or deteriorated.

Are ChatGPT citations accurate?

Some citations can be useful and well matched; others can be incomplete, wrong or outdated. Judge the specific answer and source instead of assuming that the product name determines reliability.

Start by distinguishing an answer that searched the web from one that generated references without observable retrieval. Then open the cited source. Asking the same assistant whether its citation is correct is not an independent check.

If a citation points to a page you cannot access, record the verification limitation. Try the publisher's record or another legitimate authoritative route. Do not treat every blocked page as a fake reference.

Are Gemini citations accurate?

Apply the same source checks, but also inspect the link type. Google explains that a link found by Gemini's double-check feature is not necessarily a source used to generate the original response.[4]

That link can help corroborate a statement, but counting it as an original citation changes what your report means. Preserve the distinction between source references, related material and later verification results.

How to verify an AI citation in five steps

1. Separate the answer into checkable claims

A paragraph may contain several claims supported by one marker. Split them before checking.

For example, “This platform supports offline work, costs less than its competitors and is the market leader” contains three different assertions. A product page might support the first, a price comparison the second, and neither the third.

2. Verify the source's identity

Open the link and check the title, publisher, author and date where applicable. For academic references, verify the bibliographic record rather than trusting a plausible-looking DOI string.

Keep the cited URL and the final destination. If the page redirects to a homepage, the original claim may no longer be directly verifiable.

3. Find the supporting passage

Search within the source for the relevant term or number, then read the surrounding context.

Ask:

  • Is this the same population, product version, location and period?
  • Is the number a total, average, estimate or maximum?
  • Is the source reporting an observation, prediction or opinion?
  • Has the answer dropped a qualification?

A topical match is weaker than direct support.

4. Check freshness and stronger sources

For a current product feature, use current official documentation. For research, distinguish a preprint from a peer-reviewed version and check for corrections. For reporting, prefer the original source over an unexplained copy when possible.

The right freshness standard depends on the claim. A mathematical definition and a software price should not receive the same update interval.

5. Record a verdict with evidence

Use a small set of labels:

  • Fully supported.
  • Partially supported.
  • Contradicted.
  • Source exists but does not support the claim.
  • Unverifiable with available access.
  • Reference not found after the documented search.

Do not merge “unverifiable” and “fabricated.” Save a short note explaining the verdict and the passage used.

A worked verification example

Suppose an AI answer says, “A study proved this content change increases sales by 40%,” and cites a paper about visibility in generated answers.

The source may be real, but the answer has changed both the outcome and the certainty. A visibility measure is not sales, and a result from one experimental setting does not prove a universal effect.

The verdict should identify the mismatch, not merely mark the URL as valid. A corrected sentence would name the actual visibility metric, experimental setting and limitation.

This is an illustrative verification exercise, not a newly observed failure by a named engine.

How to audit citations across a set of answers

Choose the sample before inspecting which engine looks best. Define the topic, dates, product modes and inclusion rules. Retain failures and unanswered questions.

For each answer, check both:

  1. Citation support: do supplied sources support the associated claims?
  2. Support coverage: which externally verifiable claims lack adequate support?

Add URL availability, source identity and freshness as separate checks. Do not average them into a single score unless the weighting is explicit and useful for your decision.

If several reviewers are involved, agree on examples first and review disagreements. For repeated prompts, retain prompt IDs so the analysis does not pretend every answer is independent.

Use this citation-audit template and study-comparability table to structure your own review.

What citation errors mean for brands

A citation to your website can still misrepresent your product. That makes accuracy review part of brand monitoring, not just an academic concern.

Prioritise material errors: unavailable features, incorrect pricing, wrong markets, obsolete policies or misleading comparisons. Identify the source of the claim where visible, correct your own information where necessary, and document the issue before seeking a legitimate correction elsewhere.

AI citation tracking helps preserve the source and answer context. It does not replace the human judgement needed to decide whether a claim is actually supported.

Frequently asked questions

Does a DOI prove that a citation is correct?

No. A DOI can establish a publication's identity, but the paper may not support the claim attached to it. Verify both the record and the relevant content.

Are paid AI tools always more accurate?

Do not infer accuracy from price alone. Evaluate the actual task, model, retrieval mode and evidence. A tool that declines an uncertain question also needs to be evaluated differently from one that always supplies an answer.

Should I stop using AI for research?

AI can help discover candidates, organise questions and summarise material. Keep source verification and consequential judgement in the workflow. For a practical discovery process, see Can ChatGPT find citations?.

Sources

  1. Liu, Zhang and Liang, Evaluating Verifiability in Generative Search Engines, EMNLP Findings 2023.
  2. Jaźwińska and Chandrasekar, Tow Center, AI Search Has a Citation Problem, 6 March 2025.
  3. Liu et al., Do AI chatbots find what experts would?, August 2026 preprint. Medical evidence-retrieval task, not clinical advice or a general accuracy benchmark.
  4. Google, Gemini sources and double-checking.