# We Asked ChatGPT the Same Questions Four Times. Only 12% of Sources Showed Up Every Time.

**URL:** https://organikpi.com/blog/geo-ai-search/chatgpt-citation-variance-study/
**Published:** 2026-08-16
**Modified:** 2026-08-20
**Author:** Daniel Shashko

> ChatGPT's citation behavior is stable in aggregate and unstable per answer, so any tracker that reports a single yes-or-no citation check is reporting a coin flip. Across 28 questions run four times in one day, 56% of cited sources appeared in exactly one run and 11.8% appeared in all four. The average overlap between two runs of the same question was 0.39. Eight of the 28 questions credited zero sources in one run and several in another. The pooled credit rate held between 22.0% and 27.9%. Across 64 questions ChatGPT surfaced 1,008 sources and credited 238, or 23.6%. Homepages were 0.4% of credited URLs.

---

> ChatGPT's citation behavior is stable in aggregate and unstable per answer, so any tracker that reports a single yes-or-no citation check is reporting a coin flip. Across 28 questions run four times in one day, 56% of cited sources appeared in exactly one run and 11.8% appeared in all four. The average overlap between two runs of the same question was 0.39. Eight of the 28 questions credited zero sources in one run and several in another. The pooled credit rate held between 22.0% and 27.9%. Across 64 questions ChatGPT surfaced 1,008 sources and credited 238, or 23.6%. Homepages were 0.4% of credited URLs.

We sent 28 questions to ChatGPT four times each in a single day. The source lists came back different almost every time. Across the four runs, 56% of the sources ChatGPT credited turned up in exactly one run, and only 12% turned up in all four. Eight of the 28 questions produced a full list of sources in one run and no sources at all in another.

The pooled number stayed flat. The share of pulled-up sources that got credited landed at 23.1%, 24.9%, 27.9% and 22.0% across the four runs. ChatGPT&#8217;s overall citation behavior is a stable, measurable property. Whether your page appears in any single answer is close to a coin flip.

That gap matters because most [AI citation tracking tools](https://organikpi.com/blog/seo-strategy/ai-search-competitive-intelligence-tools/) report the second thing and present it as the first.

## What we ran

64 questions, written to look like what people actually type into ChatGPT: comparisons, product research, how-tos, definitions, recency questions, local pricing, and statistics. Eight of the 64 were non-search controls. Every prompt went through Bright Data&#8217;s ChatGPT scraper on 16 August 2026, US location, on the gpt-5-6 default model.

28 of those questions, four from each of the seven search categories, then went through three more identical runs the same day. The first run counts as a fourth replicate, so the repeat test rests on 112 answers. The full set surfaced 1,008 sources.

One number this setup cannot produce is how often ChatGPT decides to search the web at all. The scraper submits through ChatGPT&#8217;s search entry point, which pushes that decision toward yes. Everything below is conditioned on a search having run, and there are no trigger rates in this post.

## The same question, four different source lists

			
				
			
		Each row is one question. Each dot is one of four runs on the same day.

The extremes are not subtle. &#8220;What changed in Google Search in 2026&#8221; credited 12 sources, 12 sources, zero sources, then 12 again. &#8220;Latest electric vehicle tax credit rules in the US&#8221; went 4, 2, 0, 0. &#8220;How to set up a custom domain on Cloudflare&#8221; went 0, 5, 6, 7.

Eight of 28 questions hit zero at least once while citing sources in another run. The average swing in how many sources an answer credited was 38% of that question&#8217;s own mean. Anything smaller than that is noise, which sets a floor on what a week-over-week change can mean.

## More than half the sources appeared in only one run

			
				
			
		204 question-and-source pairs across 28 questions and four runs.

Pooling every question with every domain it ever cited gives 204 pairs. 114 of them (55.9%) showed up in a single run. 24 of them (11.8%) showed up in all four. Comparing any two runs of the same question, the source lists overlapped by 0.39 on average, with a median of 0.33.


    Key Claim
    A single ChatGPT sample sees a given cited domain in about one run out of four. Presence in one answer is not evidence of presence, and absence in one answer is not evidence of absence.


            
                            OrganiKPI, 28 prompts x 4 runs, 16 August 2026                    
    

Read that against how visibility is usually reported. A tool that samples a prompt once a week and tells you that you were cited is describing one draw out of many. The same is true when it tells you a competitor took your spot. This is the measurement problem sitting underneath every [share of voice](https://organikpi.com/blog/seo-strategy/ai-search-brand-share-of-voice/) number in the category.

## The site-wide rate holds while the individual answer moves

			
				
			
		Pooling 28 questions gives a steady rate. Reading them one at a time does not.

Both panels come from the same 112 answers. Pool them and the credit rate sits in a five-point band across four runs. Split them by question and the overlap between runs collapses to a third.

This tells you which questions to ask of your own data. &#8220;What share of the answers in my category cite anyone at all&#8221; is answerable. &#8220;Did we hold the citation on this prompt since last Tuesday&#8221; mostly is not, at one sample per week. The fix is more samples per prompt, and it is the same logic that makes [citation velocity](https://organikpi.com/blog/seo-strategy/citation-velocity-measurement-framework/) a more honest metric than a citation checkbox.

## ChatGPT credits about a quarter of what it pulls up

			
				
			
		Sources surfaced against sources credited, across all 64 questions.

Across the 64 questions, ChatGPT surfaced 1,008 sources and credited 238 of them. That is 23.6%, or roughly four candidates for every link that made it into an answer. The rate held between 20% and 30% in every question category we ran.

Getting surfaced and getting credited are separate outcomes, and the second one is where the traffic is. Ranking for the search ChatGPT writes for itself only puts you in the candidate pool.

Those self-written searches are worth looking at directly. Across 48 answers ChatGPT wrote 81 searches of its own, a median of 2 per question and up to 5. 52% of answers ran more than one. Your page competes against whatever ChatGPT decided to type, which is why [prompt research](https://organikpi.com/blog/seo-strategy/prompt-research-vs-keyword-research/) and keyword research pull apart.

Credit is also spread thin. The 238 credits went to 173 different domains. The most-cited domain took 3.8% of them. The top ten combined took 18.5%. 78% of the domains that earned a credit earned exactly one. There is no small club of authorities in this sample, which is a different picture from the concentration we found when we [decoded 42,971 Google AI Mode citations](https://organikpi.com/blog/geo-ai-search/decoded-42971-ai-citations-google-research/).

## What separates a credited source from a dropped one

Pooling all four runs gives 639 credited sources and 2,003 dropped ones. Two things separate them, and both are smaller than they are usually made out to be.

SignalCreditedDroppedGapPublished 2026 or later, of sources showing a date69.6%62.0%7.5 pointsURL folder depth, median321 folderSample size6392,003Recency is measured on the 230 credited and 1,127 dropped sources where ChatGPT surfaced a publication date.

Recency helps. A 7.5 point edge clears statistical significance at z = 2.16, and it is worth keeping your dates real. Most [content freshness](https://organikpi.com/blog/content-strategy/content-freshness-recency-bias/) advice treats it as a much bigger lever than 7.5 points, and republishing an old page under a new date buys nothing at all.

Depth is the more useful signal. Credited URLs sit a folder deeper than dropped ones. Combined with the near-total absence of homepages below, the pattern says a page built to answer one question beats a hub page that covers ten, which is the case for [cluster pages over pillar pages](https://organikpi.com/blog/content-strategy/pillar-cluster-content-geo-strategy/) when the goal is a citation.

One signal looked far stronger than either and turned out to be worthless. Credited sources sat at a median position of 3 in the source list against 16 for dropped ones. In 134 of 134 answers, the credited set occupied positions 1 through k exactly. Position is assigned after the citation decision, so it predicts nothing. It is the kind of number that reaches a slide deck intact.

## Where our numbers land on Similarweb&#8217;s

			
				
			
		Our 238 cited URLs against Similarweb&#8217;s 2026 Generative AI Landscape measurement.

Aleyda Solis of Orainti, analyzing the data in Similarweb&#8217;s 2026 Generative AI Landscape report, put 65% of ChatGPT-cited URLs two or three folders deep and 41.7% at exactly two. We got 63.4% and 39.1% on a sample thousands of times smaller. Landing inside two points on an independent apparatus is a good sign for both measurements.


    1 of 238
            credited URLs that were a homepage
                
                            OrganiKPI, 64 ChatGPT prompts, 16 August 2026                    
    

Homepages were 0.4% of credited URLs. Set that next to Solis&#8217;s finding, in the same report, that 58.8% of AI referral traffic lands on homepages and the two describe opposite ends of the same session: ChatGPT cites the deep page, then the user goes looking for the brand. That is a reason to keep your [GA4 referral attribution](https://organikpi.com/blog/technical-seo/ga4-ai-search-referral-attribution/) pointed at the whole path rather than the landing page alone.

## What to do about it

### 1. Report citation presence as a frequency, never as a yes or no

&#8220;Cited on 6 of 20 samples this month&#8221; is a measurement. &#8220;Cited&#8221; is a draw from a distribution. Rewrite your internal reporting first, before you change anything about the content.

### 2. Ask any tracker how many samples it takes per prompt per period

This is the single question that separates a usable tool from a decorative one, and it is rarely on the pricing page. Ask it during the demo of every tool on your [AI visibility tool](https://organikpi.com/blog/geo-ai-search/best-ai-visibility-tools/) shortlist. One sample per prompt per week cannot support a week-over-week chart.

### 3. Ignore movements smaller than a third

The average question&#8217;s credited-source count swung by 38% of its own mean between identical runs. Treat that as your noise floor. A 20% drop in citations week over week is not a finding, and chasing it will send you into an [AI citation recovery](https://organikpi.com/blog/geo-ai-search/ai-citation-recovery-after-update/) project you do not need.

### 4. Track the pooled rate across many prompts, not presence on a few

The stable quantity in this study was the aggregate. Thirty prompts sampled once each will tell you more about your position than three prompts sampled ten times. Build the prompt set wide, then watch the share, which is what a working [GEO KPI framework](https://organikpi.com/blog/seo-strategy/geo-kpi-measurement-framework/) measures.

### 5. Give each real question its own page

Credited URLs sat a folder deeper than dropped ones and homepages were 0.4% of credits. A page that answers one question completely is the shape that gets cited. This is also the cheapest fix on the list, because it usually means splitting a page you already have rather than writing a new one, and it doubles as [content pruning](https://organikpi.com/blog/content-strategy/ai-search-content-pruning/) work.

### 6. Keep dates honest, and size the expectation

Recency bought 7.5 points among sources that showed a date. Keep publication dates accurate and update pages when the facts change. Do not build a content plan around it.

### 7. Optimize for the search ChatGPT writes, not the one your user types

Half the answers ran more than one self-written search. Those queries are visible in the response data, and they are a better keyword list than anything a keyword tool will give you for [ranking in ChatGPT search](https://organikpi.com/blog/geo-ai-search/how-to-rank-in-chatgpt-search/).

## What this changes

ChatGPT&#8217;s citation behavior is measurable in aggregate and unstable per answer. Both statements are true at the same time, and the category has mostly been selling the second one as though it were the first.

AI visibility is still measurable. The unit is a rate over many samples, the same way nobody judges a Google ranking from one incognito search. Set your prompt set wide, sample it repeatedly, report frequencies, and ignore anything smaller than a third. The tools in the [GEO tooling category](https://organikpi.com/blog/geo-ai-search/best-geo-tools/) can do this. Most of them are just not configured to, and the ones that are will tell you their sample count when you ask.

The raw records, the prompt set, and the analysis scripts behind every number here are available on request. If you want the wider context on how AI engines pick sources, our [State of AI Search](https://organikpi.com/blog/geo-ai-search/state-of-ai-search/) roundup and our [LLM SEO guide](https://organikpi.com/blog/geo-ai-search/llm-seo/) cover the mechanics, and [what AI visibility actually means](https://organikpi.com/blog/geo-ai-search/what-is-ai-visibility/) covers the scoring.

## Frequently Asked Questions

### How much do ChatGPT's citations change between identical runs?

A lot. Across 28 questions run four times in a single day, 56% of the sources ChatGPT credited appeared in only one of the four runs, and just 11.8% appeared in all four. The average overlap between any two runs of the same question was 0.39, with a median of 0.33.

### Does this mean AI visibility tracking does not work?

It works when the unit of measurement is a rate across many samples. The pooled credit rate held between 22.0% and 27.9% across four runs, so aggregate measurement is reliable. A single sample of a single prompt is not, because the same question credited zero sources in one run and a dozen in another.

### What percentage of the sources it looks at does ChatGPT actually cite?

Across 64 questions, ChatGPT surfaced 1,008 sources and credited 238 of them inside the answers. That is 23.6%, or roughly one credit for every four candidates. The rate stayed between 20% and 30% in every question category tested.

### Does publication date affect whether ChatGPT cites a source?

Slightly. Among sources where ChatGPT surfaced a publication date, 69.6% of credited sources were published in 2026 or later against 62.0% of dropped ones. That is a 7.5 point edge. It is real but small, and it does not support republishing old pages under new dates.

### Does ChatGPT cite homepages or deep pages?

Deep pages, almost exclusively. Only 1 of 238 credited URLs was a homepage. 63.4% sat two or three folders deep, which lands within two points of Similarweb's 65% measurement on a far larger sample. Credited URLs sat a median of one folder deeper than dropped ones.

### How many samples per prompt does an AI visibility tracker need?

More than one per reporting period. The number of sources an answer credited swung by an average of 38% of that question's own mean between identical runs, so any week-over-week movement smaller than roughly a third sits inside the noise floor. Ask any vendor how many samples they take per prompt per period.

### How many searches does ChatGPT run for a single question?

A median of two, and up to five. Across 48 answers ChatGPT wrote 81 searches of its own, and 52% of answers ran more than one. Those self-written queries, not the user's original wording, are what your page competes against.

