# ChatGPT Writes Its Own Search Query First. We Captured 219 of Them

**URL:** https://organikpi.com/blog/geo-ai-search/chatgpt-query-fanout-study/
**Published:** 2026-09-15
**Modified:** 2026-09-16
**Author:** Daniel Shashko

> ChatGPT writes its own search query before it looks anything up. 207 of 209 captured queries have no Keyword Planner data. 70 are too long for Google Ads to accept. A third name a brand the person never mentioned. A quarter go straight to a vendor's own site.

---

> ChatGPT writes its own search query before it looks anything up. 207 of 209 captured queries have no Keyword Planner data. 70 are too long for Google Ads to accept. A third name a brand the person never mentioned. A quarter go straight to a vendor's own site.

When ChatGPT decides to search the web, it does not send your question to the search engine. It writes its own query first. We captured 219 of those queries in August 2026, and 207 of the 209 unique ones return no data at all in Google&#8217;s Keyword Planner.

That layer decides which pages get pulled in to answer a question about your category, and no keyword tool can see it. On buyer questions it is worse than invisible. The query that goes out to find evidence already carries a shortlist of vendors, named by the model before any search ran. 36.1% of the queries we captured name a company, product or publication the person never typed.

## One question, five searches

Someone asked for the best project management software for a 10 person agency. ChatGPT ran five searches. The first one named Asana, ClickUp, Monday, Teamwork and Basecamp, none of which appear in the question. The other four each checked one of those vendors against its own site, using the word &#8220;official&#8221; to steer away from third-party pages.

			
				
			
		One question, five searches. ChatGPT named the shortlist inside the query it sent out to find evidence for it.

Across all 121 captures, one question produced 1.81 searches on average. Fifty-one produced a single query, 54 produced two, and 16 produced three or more. The longest fan-out for one question was six.

## The queries do not exist in any keyword tool

We put all 209 unique fan-out queries through Google Ads. Two came back with a search volume. The other 207 did not exist as far as the Keyword Planner is concerned, and 70 of those were rejected outright rather than returned with a zero. Google Ads caps a keyword at [80 characters and no more than 10 words](https://experienceleague.adobe.com/en/docs/advertising/search-social-commerce/campaign-management/management/campaigns/keywords/keyword-settings-by-network/keyword-settings-google), and a third of what ChatGPT types is longer than that.

The same test on the 53 prompts people actually typed returns 14 with volume, so 73.6% were invisible. Conversational questions are already hard to research. The machine&#8217;s rewrite of them is a different category of hard.

Both exceptions are worth reading before quoting the headline. how to remove a stripped screw returns 33,100 a month, and it is the one case in the whole set where ChatGPT did not rewrite the question at all. site:zoho.com crm pricing returns 2,400, but only because Google Ads strips the operator and prices the words that are left. Neither is evidence that fan-out queries are researchable.

			
				
			
		Google Keyword Planner, United States, pulled September sixteenth. The rejected share is over the keyword length cap.

## It names the vendors before it looks them up

79 of the 219 queries, 36.1%, carry a company, product or publication that is absent from the user&#8217;s prompt. The median injection is two names and the maximum is eight. Ask for password managers for a family and the model writes 1Password, Bitwarden and Dashlane into the query, then runs three more searches restricted to those three domains.

The split by question type is sharper. On shortlist-shaped prompts, the &#8220;best X for Y&#8221; and &#8220;is X worth it&#8221; questions where a buying decision is at stake, 18 of 27 prompts had a name injected. On everything else it was 25 of 94. Those are small counts and we are reporting them as counts for that reason, but the direction is not subtle.

The sequence matters more than the percentage. The candidate set is in the query that goes out to gather evidence, which means it came from the model&#8217;s memory rather than from the search. Retrieval is being used to check a shortlist, not to build one. That is a different problem from the one most [answer engine optimization](https://organikpi.com/blog/geo-ai-search/what-is-aeo-answer-engine-optimization/) advice is aimed at, and it sits upstream of the source instability we measured when we [asked ChatGPT the same questions four times](https://organikpi.com/blog/geo-ai-search/chatgpt-citation-variance-study/).

## A quarter of the searches go to a vendor&#8217;s own site

24.2% of the queries aim at a primary source. 11.0% use a site: operator to restrict the search to one domain. 13.2% carry the word &#8220;official&#8221;. 15.1% contain the word &#8220;pricing&#8221;. The pattern repeats across categories: name the vendor, then go to the vendor to check what it costs and what it does.

Your pricing page is a retrieval target. So are your docs and your feature pages. If the number a buyer needs is behind a form, the model runs that search, finds nothing usable, and takes a third-party figure instead. Braze&#8217;s own [AWS Marketplace listing showed $75,000 while its pricing page named no number at all](https://organikpi.com/blog/geo-ai-search/contact-sales-pricing-already-public/).

62.6% of the queries also carry a year, usually &#8220;2026&#8221;, appended to questions that named no date. A page that does not state when its facts were true is competing against a query that explicitly asks for this year.

			
				
			
		What the model adds to your question before it searches. Share of all captured fan-out queries.

## What this changes

Keyword research cannot reach this layer, which is the practical version of the argument for [prompt research over keyword research](https://organikpi.com/blog/seo-strategy/prompt-research-vs-keyword-research/). You cannot target a query with no volume, and you cannot target one Google Ads refuses to accept. What you can do is work the two ends the fan-out actually touches.

- **Get into the candidate set.** The shortlist comes from the model&#8217;s memory, which is built from mentions across the wider web rather than from your own pages. That is the case for the off-site work in our [LLM SEO guide](https://organikpi.com/blog/geo-ai-search/llm-seo/) and for tracking [share of voice](https://organikpi.com/blog/seo-strategy/ai-search-brand-share-of-voice/) rather than rankings.
- **Win the verification search.** Once you are named, the next query goes to your domain with &#8220;official&#8221; or site: attached. Put prices, limits and specifications in visible text on the page, which is also what the [Perplexity retrieval playbook](https://organikpi.com/blog/geo-ai-search/perplexity-citation-strategy/) depends on.
- **Date your facts.** Nearly two thirds of these queries ask for a specific year.
- **Write extractable sentences.** The retrieved page still has to yield a usable claim, which is what [atomic sentence structure](https://organikpi.com/blog/content-strategy/atomic-sentence-seo-ai-citations/) is for and what our [analysis of 42,971 AI citations](https://organikpi.com/blog/geo-ai-search/decoded-42971-ai-citations-google-research/) found Google rewarding.
- **Cover the sub-questions, not just the head term.** One question becomes almost two searches, so [pillar and cluster architecture](https://organikpi.com/blog/content-strategy/pillar-cluster-content-geo-strategy/) and [gap analysis](https://organikpi.com/blog/content-strategy/content-gap-analysis-ai-search-era/) matter more than a single page that ranks.

None of this is specific to one engine&#8217;s interface. It is how a model with a search tool behaves. Ask [four engines the same question and four different shortlists come back](https://organikpi.com/blog/geo-ai-search/ai-engines-different-recommendations/), which is what separate memories producing separate candidate sets looks like from the outside.

The tactics in our [ChatGPT citation playbook](https://organikpi.com/blog/geo-ai-search/how-to-rank-in-chatgpt-search/) still apply, and so does the retrieval work behind [Google AI Mode](https://organikpi.com/blog/geo-ai-search/google-ai-mode-how-it-works/). This just says which step they are aimed at.

## Sample, and what it does not show

The captures ran on 16 August 2026 against ChatGPT, logged out, from a US exit. 121 of them fired a web search and returned the query the model wrote. Keyword Planner status was pulled on 16 September 2026 from Google Ads, United States, English. Entity matching runs on word boundaries against a hand-built list, and generic tokens inside product names are excluded, so 36.1% is a floor. The queries themselves are not new: they ship as a column in the dataset behind our [citation variance study](https://organikpi.com/blog/geo-ai-search/chatgpt-citation-variance-study/), published in August. What is new here is the analysis of them.

Four limits are worth stating plainly. This is one engine; Google AI Mode exposes no equivalent field. It is logged out and US-only, so there is no personalization arm. 121 captures describe the shape of a fan-out query well and are too thin for a stable per-category rate. And the capture cannot be repeated: the field that carries these queries returned data on 16 August and returns null as of 15 September, on the same code path, with the search still firing. We have reported that to the vendor. The dataset below is what exists.

A related question we cannot answer from this data is whether the injected shortlist is stable. Across three repeat runs of the same prompts the overlap was low, but only 15 prompts appeared in all three runs, which is too few to publish a rate. We have left it out rather than dress it up. For what repetition does to the sources behind an answer, the [Reddit citation study](https://organikpi.com/blog/geo-ai-search/chatgpt-reddit-citation-collapse/) covers the same ground with a workable sample.

## Frequently Asked Questions

### What is query fan-out in ChatGPT?

Query fan-out is the set of searches a model writes for itself before it answers. Rather than passing your question to a search engine, ChatGPT composes one or more of its own queries, runs them, and answers from what comes back. In our captures one question produced 1.81 searches on average, and the longest fan-out was six.

### Can I find ChatGPT's fan-out queries in a keyword tool?

No. We put 209 unique fan-out queries through Google Ads and 207 returned no search volume. Seventy of them were rejected outright because Google Ads accepts a keyword of at most 80 characters and 10 words, and a third of the queries were longer than that.

### Does ChatGPT decide which brands to recommend before it searches?

The brand names appear in the query that goes out to gather evidence, which means they came from the model rather than from the search results. 36.1% of the captured queries named a company, product or publication absent from the user's prompt. On shortlist-shaped buying questions it was 18 of 27 prompts.

### Why does ChatGPT search a vendor's own website?

It checks named vendors against their primary source. 11.0% of the queries used a site: operator to restrict the search to one domain and 13.2% carried the word official, with 15.1% asking about pricing. That makes a public pricing page, docs and feature pages the pages being retrieved.

### How do you optimize for query fan-out?

Work both ends of it. Earn mentions across the wider web so the model names you in the first place, then make sure the verification search finds real numbers in visible text on your own site. Date your facts, since 62.6% of the queries appended a year the user never mentioned.

