TL;DR
- About 85% of AI citations back at least part of the claim they are attached to, and 15% back none of it.
- Missing feature details cause most partial verdicts, and prices are second at 16%.
- ChatGPT has the lowest unsupported rate at 7.3%, Gemini the highest at 22.8%.
- Vendor-owned pages fail verification at 23%, nearly double the 13% rate for third-party review sites.
- The dataset covers 2,664 judgeable claims from 160 product-comparison queries across ChatGPT, Gemini, and Copilot.
AI engines attach citations to make their answers look grounded. We checked 2,664 of those claims against the pages they actually link to. 85% of citations back at least part of the claim. 15% back none of it.
The question behind this study is simple: when ChatGPT, Gemini, or Copilot cites a page, does that page say what the engine claims it says? We ran 160 product-comparison queries through all three engines, extracted every citation-to-claim pair, fetched the cited pages, and judged each one.
The headline numbers

Of 2,664 judgeable citation-claim pairs across ChatGPT, Gemini, and Copilot:
- 1,409 (53%) were fully supported: the cited page states the substance of the claim, with names, numbers, and qualities matching.
- 855 (32%) were partially supported: the page covers the topic and confirms part of the claim, but a material detail (a price, a plan name, a comparison) is missing or different.
- 400 (15%) had no support: the cited page does not say what the engine attributed to it, or contradicts it outright.

The split between full and partial support is softer than the gap between “some backing” and “none.” A blind second labelling of 60 random rows agreed on “not supported vs. everything else” 92% of the time (Cohen’s kappa 0.55 overall, with most disagreements on the supported-vs-partial boundary). The first judge was slightly harsher on not-supported verdicts (8 of 60 vs. 5 of 60), so the 15% figure may be a slight overcount.
ChatGPT is the most accurate, Gemini the least
| Engine | Claims judged | Fully supported | Partially supported | Not supported |
|---|---|---|---|---|
| ChatGPT | 915 | 572 (62.5%) | 276 (30.2%) | 67 (7.3%) |
| Copilot | 840 | 372 (44.3%) | 342 (40.7%) | 126 (15.0%) |
| Gemini | 909 | 465 (51.2%) | 237 (26.1%) | 207 (22.8%) |
ChatGPT’s unsupported rate is 7.3%, roughly a third of Gemini’s 22.8%. Copilot sits in the middle at 15%. The gap holds after adjusting for the inter-rater softness on the boundary: even on the binary question of “does the page back the claim at all,” ChatGPT outperforms.
Copilot has the highest partial rate (40.7%), meaning it often cites a page that covers the right topic but misses a specific detail the engine added. Gemini has the lowest partial rate (26.1%), but that is because more of its misses land in the unsupported bucket instead.
For context, our comparison of 1,920 AI answers across the same queries found that engines disagree with each other on brand recommendations roughly half the time. The gap in AI visibility between engines extends to citation quality. Citation accuracy adds another layer: even when two engines name the same product, the pages they cite to justify that choice may say different things.
Why partial citations fall short
Most partial verdicts come from a missing detail. Of 855 partially supported claims, 64% include a feature or spec the cited page never mentions. Prices are the second cause at 16%. Ranking language such as “best overall” accounts for 10%, and plan or tier names for 9%.
A typical price miss: ChatGPT said Intercom starts at $29/seat/month annually, with its Fin AI agent charged per outcome, and cited Intercom’s pricing page. The page confirms the per-outcome Fin pricing. It never states the $29/seat figure.
Our study of hidden pricing in AI answers covers the same gap from the vendor side.

For brands, the practical step is to state prices, plan names and key features in plain text on the pages you want cited. A citation cannot back a detail its page does not contain.
Vendor pages are cited often and verified less
We classified each cited URL by whether it belongs to the vendor named in the claim. Of 494 citations pointing to a vendor’s own site, 23% were not supported. For third-party pages (review sites, comparison articles), the rate was 13%.
Vendor pages are more likely to fail verification because the engine attributes a specific claim (a price, a feature comparison, a “best for” ranking) and links it to a page that exists to sell, not to confirm a factual comparison. A product marketing page will describe its own features but will not say “we are cheaper than Competitor X,” even when the engine’s answer says exactly that.
Third-party review sites, by contrast, tend to make the kind of comparative statements that engines attribute: “best for small teams,” “strongest reporting.” When an engine cites a review site, the language in the answer often tracks what the reviewer wrote. This fits the pattern from our Reddit citation research: engines tend to cite pages that make explicit comparative statements.
What an unsupported citation looks like
Three real cases from the dataset:
- Wrong product on the page. Gemini cited a “best accounting software” article from TGG Accounting and attributed a recommendation for Sage 50. The cited page lists QuickBooks, Xero, Zoho Books, FreshBooks, and Wave. Sage 50 does not appear anywhere on it.
- Right vendor, wrong page. ChatGPT cited a Bitwarden marketing page and stated “$4/user/month.” The linked page is generic copy with no pricing. The price may be correct, but the citation links to a page that does not contain it.
- Composite claim, single source. ChatGPT compared Five9 and Aircall in one sentence, then cited a single review page. The page confirmed Five9’s 50-seat minimum but only hinted at Aircall’s positioning through its lower price. Half the claim checks out, and the other half rests on inference.
128 of 160 queries produced at least one unsupported citation across the three engines.
Not every citation can be checked
The three engines produced 3,970 total citations across 160 queries. Of those, 3,059 (77%) could be mapped to a specific claim-sentence pair. The rest either pointed to pages without a clear sentence-level attribution or appeared as generic “learn more” links.
Of the mapped citations, 354 (12%) pointed to pages that were unreachable (paywalls, login walls, geographic blocks, or pages that returned errors during collection). Those are excluded from the accuracy numbers above.
| Engine | Total citations | Mapped to a claim | Unreachable |
|---|---|---|---|
| ChatGPT | 1,582 | 1,109 (70.1%) | 158 (14.4%) |
| Gemini | 1,109 | 1,014 (91.4%) | 104 (10.3%) |
| Copilot | 1,279 | 936 (73.2%) | 92 (9.8%) |
Gemini maps 91% of its citations to specific sentences, the highest rate. ChatGPT and Copilot both sit around 70-73%. This difference is partly structural: Gemini tends to attach inline citations to individual claims, while ChatGPT groups citations at the paragraph level, making sentence-level attribution harder to extract.
How the study was run
We used the same 160 product-comparison queries from our AI Overview citation study and self-promotional listicle analysis. All queries follow the “best [product category]” pattern. Responses were collected from ChatGPT, Gemini, and Copilot in October 2026, using Bright Data’s AI scraper endpoints. Google AI Mode was attempted but dropped after the collection pipeline stalled. Perplexity was excluded because of a sign-up wall on the scraper endpoint.
Each response was parsed to extract citation-claim pairs: the sentence a citation is attached to, and the URL it points to. The cited pages were fetched and their text extracted. Each pair was then judged by an LLM judge with a verbatim-quote check (the judge must return an exact substring from the page as evidence, or the verdict is downgraded). A blind second labelling of 60 random rows (20 per engine) produced 73% exact agreement and 92% agreement on the binary “supported vs. not” question.
The query set covers only product-comparison queries (“best CRM,” “best password manager”). Results may differ for informational, navigational, or opinion queries. All three engines were queried once per query; our variance study found that ChatGPT’s citations change substantially across repeated runs, so a second run would produce different individual verdicts. The aggregate rates should be stable. The broader adoption data on AI search gives context on how many people rely on these citations daily.
What this means for brands and publishers
If you publish content that AI engines cite, 15% of claims attributed to your page may not actually appear on it. That is a reputational risk you cannot control directly, but you can reduce it.
- Make factual claims explicit and crawlable. Prices, feature comparisons, and “best for” statements should appear as visible text, not behind tabs, accordions, or JavaScript rendering. If the engine can read the text, it is more likely to quote it correctly.
- Publish stable pages. Pricing pages that change monthly create a rolling window of stale citations. Our data shows pricing is the single largest category of partial citation failure.
- Check what engines say about you. Run your brand queries through ChatGPT, Gemini, and Copilot. Click the citations. If the linked page does not say what the engine claims, that is something a visibility audit can catch before a prospect does.
- Watch for composite claims. Engines often compare two or three products in a single sentence and attach one citation. That citation may back one product’s claim but say nothing about the others. If your product is the one without backing, the citation is cosmetic. This matters most for answer engine optimization strategies that track citation counts without checking citation quality.
For publishers who write the comparison articles that engines cite, the same logic applies in reverse. If your review says “Product X starts at $49/month” and that price changes, the engine will keep citing your old number. Freshness matters for citation accuracy as much as it matters for ranking.
The broader pattern from our LLM citation research: AI engines cite sources to signal credibility, and the citations are mostly correct. But “mostly” means roughly one in six claims has no backing at all on the page it links to. That gap is worth knowing about, whether you are the one being cited or the one reading the answer.
Limits
- All 160 queries are product comparisons (“best X”). Accuracy rates for factual, how-to, or opinion queries are likely different.
- Single-run collection. Citation variance means individual verdicts would change on a second run; aggregate rates should be stable.
- The judge is an LLM, not a human. Inter-rater agreement on the supported-vs-not boundary (92%) is high, but the full-vs-partial boundary (78%) is softer. Some claims classified as not-supported may be partially supported by a stricter reader.
- Google AI Mode and Perplexity are not included. AI Mode‘s collection stalled; Perplexity requires a sign-up wall the scraper could not bypass.
- Cited pages were fetched once on the same day as the AI responses. Pages that changed between the engine’s last crawl and our fetch may show a false mismatch.