GEO & AI Search

We Asked ChatGPT the Same Questions Four Times. Only 12% of Sources Showed Up Every Time.

Updated 7 min read Daniel Shashko
We Asked ChatGPT the Same Questions Four Times. Only 12% of Sources Showed Up Every Time.
TL;DR
  • ChatGPT's citation behavior is stable in aggregate and unstable per answer, so any tracker that reports a single yes-or-no citation check is reporting a coin flip.
  • Across 28 questions run four times in one day, 56% of cited sources appeared in exactly one run and 11.8% appeared in all four.
  • The average overlap between two runs of the same question was 0.39.
  • Eight of the 28 questions credited zero sources in one run and several in another.
  • The pooled credit rate held between 22.0% and 27.9%.
  • Across 64 questions ChatGPT surfaced 1,008 sources and credited 238, or 23.6%.
  • Homepages were 0.4% of credited URLs.

We sent 28 questions to ChatGPT four times each in a single day. The source lists came back different almost every time. Across the four runs, 56% of the sources ChatGPT credited turned up in exactly one run, and only 12% turned up in all four. Eight of the 28 questions produced a full list of sources in one run and no sources at all in another.

The pooled number stayed flat. The share of pulled-up sources that got credited landed at 23.1%, 24.9%, 27.9% and 22.0% across the four runs. ChatGPT’s overall citation behavior is a stable, measurable property. Whether your page appears in any single answer is close to a coin flip.

That gap matters because most AI citation tracking tools report the second thing and present it as the first.

What we ran

64 questions, written to look like what people actually type into ChatGPT: comparisons, product research, how-tos, definitions, recency questions, local pricing, and statistics. Eight of the 64 were non-search controls. Every prompt went through Bright Data’s ChatGPT scraper on 16 August 2026, US location, on the gpt-5-6 default model.

28 of those questions, four from each of the seven search categories, then went through three more identical runs the same day. The first run counts as a fourth replicate, so the repeat test rests on 112 answers. The full set surfaced 1,008 sources.

One number this setup cannot produce is how often ChatGPT decides to search the web at all. The scraper submits through ChatGPT’s search entry point, which pushes that decision toward yes. Everything below is conditioned on a search having run, and there are no trigger rates in this post.

The same question, four different source lists

The extremes are not subtle. “What changed in Google Search in 2026” credited 12 sources, 12 sources, zero sources, then 12 again. “Latest electric vehicle tax credit rules in the US” went 4, 2, 0, 0. “How to set up a custom domain on Cloudflare” went 0, 5, 6, 7.

Eight of 28 questions hit zero at least once while citing sources in another run. The average swing in how many sources an answer credited was 38% of that question’s own mean. Anything smaller than that is noise, which sets a floor on what a week-over-week change can mean.

More than half the sources appeared in only one run

Pooling every question with every domain it ever cited gives 204 pairs. 114 of them (55.9%) showed up in a single run. 24 of them (11.8%) showed up in all four. Comparing any two runs of the same question, the source lists overlapped by 0.39 on average, with a median of 0.33.

Read that against how visibility is usually reported. A tool that samples a prompt once a week and tells you that you were cited is describing one draw out of many. The same is true when it tells you a competitor took your spot. This is the measurement problem sitting underneath every share of voice number in the category.

The site-wide rate holds while the individual answer moves

Both panels come from the same 112 answers. Pool them and the credit rate sits in a five-point band across four runs. Split them by question and the overlap between runs collapses to a third.

This tells you which questions to ask of your own data. “What share of the answers in my category cite anyone at all” is answerable. “Did we hold the citation on this prompt since last Tuesday” mostly is not, at one sample per week. The fix is more samples per prompt, and it is the same logic that makes citation velocity a more honest metric than a citation checkbox.

ChatGPT credits about a quarter of what it pulls up

Across the 64 questions, ChatGPT surfaced 1,008 sources and credited 238 of them. That is 23.6%, or roughly four candidates for every link that made it into an answer. The rate held between 20% and 30% in every question category we ran.

Getting surfaced and getting credited are separate outcomes, and the second one is where the traffic is. Ranking for the search ChatGPT writes for itself only puts you in the candidate pool.

Those self-written searches are worth looking at directly. Across 48 answers ChatGPT wrote 81 searches of its own, a median of 2 per question and up to 5. 52% of answers ran more than one. Your page competes against whatever ChatGPT decided to type, which is why prompt research and keyword research pull apart.

Credit is also spread thin. The 238 credits went to 173 different domains. The most-cited domain took 3.8% of them. The top ten combined took 18.5%. 78% of the domains that earned a credit earned exactly one. There is no small club of authorities in this sample, which is a different picture from the concentration we found when we decoded 42,971 Google AI Mode citations.

What separates a credited source from a dropped one

Pooling all four runs gives 639 credited sources and 2,003 dropped ones. Two things separate them, and both are smaller than they are usually made out to be.

SignalCreditedDroppedGap
Published 2026 or later, of sources showing a date69.6%62.0%7.5 points
URL folder depth, median321 folder
Sample size6392,003
Recency is measured on the 230 credited and 1,127 dropped sources where ChatGPT surfaced a publication date.

Recency helps. A 7.5 point edge clears statistical significance at z = 2.16, and it is worth keeping your dates real. Most content freshness advice treats it as a much bigger lever than 7.5 points, and republishing an old page under a new date buys nothing at all.

Depth is the more useful signal. Credited URLs sit a folder deeper than dropped ones. Combined with the near-total absence of homepages below, the pattern says a page built to answer one question beats a hub page that covers ten, which is the case for cluster pages over pillar pages when the goal is a citation.

One signal looked far stronger than either and turned out to be worthless. Credited sources sat at a median position of 3 in the source list against 16 for dropped ones. In 134 of 134 answers, the credited set occupied positions 1 through k exactly. Position is assigned after the citation decision, so it predicts nothing. It is the kind of number that reaches a slide deck intact.

Where our numbers land on Similarweb’s

Aleyda Solis of Orainti, analyzing the data in Similarweb’s 2026 Generative AI Landscape report, put 65% of ChatGPT-cited URLs two or three folders deep and 41.7% at exactly two. We got 63.4% and 39.1% on a sample thousands of times smaller. Landing inside two points on an independent apparatus is a good sign for both measurements.

1 of 238 credited URLs that were a homepage OrganiKPI, 64 ChatGPT prompts, 16 August 2026

Homepages were 0.4% of credited URLs. Set that next to Solis’s finding, in the same report, that 58.8% of AI referral traffic lands on homepages and the two describe opposite ends of the same session: ChatGPT cites the deep page, then the user goes looking for the brand. That is a reason to keep your GA4 referral attribution pointed at the whole path rather than the landing page alone.

What to do about it

1. Report citation presence as a frequency, never as a yes or no

“Cited on 6 of 20 samples this month” is a measurement. “Cited” is a draw from a distribution. Rewrite your internal reporting first, before you change anything about the content.

2. Ask any tracker how many samples it takes per prompt per period

This is the single question that separates a usable tool from a decorative one, and it is rarely on the pricing page. Ask it during the demo of every tool on your AI visibility tool shortlist. One sample per prompt per week cannot support a week-over-week chart.

3. Ignore movements smaller than a third

The average question’s credited-source count swung by 38% of its own mean between identical runs. Treat that as your noise floor. A 20% drop in citations week over week is not a finding, and chasing it will send you into an AI citation recovery project you do not need.

4. Track the pooled rate across many prompts, not presence on a few

The stable quantity in this study was the aggregate. Thirty prompts sampled once each will tell you more about your position than three prompts sampled ten times. Build the prompt set wide, then watch the share, which is what a working GEO KPI framework measures.

5. Give each real question its own page

Credited URLs sat a folder deeper than dropped ones and homepages were 0.4% of credits. A page that answers one question completely is the shape that gets cited. This is also the cheapest fix on the list, because it usually means splitting a page you already have rather than writing a new one, and it doubles as content pruning work.

6. Keep dates honest, and size the expectation

Recency bought 7.5 points among sources that showed a date. Keep publication dates accurate and update pages when the facts change. Do not build a content plan around it.

7. Optimize for the search ChatGPT writes, not the one your user types

Half the answers ran more than one self-written search. Those queries are visible in the response data, and they are a better keyword list than anything a keyword tool will give you for ranking in ChatGPT search.

What this changes

ChatGPT’s citation behavior is measurable in aggregate and unstable per answer. Both statements are true at the same time, and the category has mostly been selling the second one as though it were the first.

AI visibility is still measurable. The unit is a rate over many samples, the same way nobody judges a Google ranking from one incognito search. Set your prompt set wide, sample it repeatedly, report frequencies, and ignore anything smaller than a third. The tools in the GEO tooling category can do this. Most of them are just not configured to, and the ones that are will tell you their sample count when you ask.

The raw records, the prompt set, and the analysis scripts behind every number here are available on request. If you want the wider context on how AI engines pick sources, our State of AI Search roundup and our LLM SEO guide cover the mechanics, and what AI visibility actually means covers the scoring.