Where Each AI Gets Its Information: Analysis of 400,000 Responses Across 7 Engines
GEO Metrics Research analyzed 400,000 responses from 7 AI engines over 14 months. The result: each engine draws from a nearly independent source universe. They share 1 in every 10.

TL;DR
After analyzing ~400,000 responses across 7 AI engines over 14 months (March 2025 – May 2026), the central finding is this: there is no "the AIs." There are seven engines that draw from nearly independent universes of evidence. Grok cites an average of 45.6 sources per response. DeepSeek cites zero — in 100% of its responses, across all 14 months. ChatGPT includes no source in 22.7% of its responses. And when two engines answer the same question, they share on average only 1 source domain out of every 10. The citation strategy that works in one engine is, in many cases, irrelevant for the next.
When an AI recommends something, there is a question almost nobody asks: what is it based on?
Does it cite a source you can open and verify, or does it speak from a self-referential void? Does its recommendation come from Wikipedia, Reddit, a niche comparison site, an arXiv paper — or nowhere in particular?
In Editions #01 and #02 of the GEO Metrics Research Lab we saw that AIs don't agree on what to recommend, and don't even describe it the same way. This edition goes to the piece underneath everything: where they get it from.
Because a recommendation without a source is an opinion with good writing.
Methodology
~400,000 responses analyzed · 7 engines that expose sources · 14 months of data · March 2025 – May 2026.
Engines analyzed: Grok, AI Mode, ChatGPT, AI Overviews, Perplexity, Copilot and DeepSeek. Only engines that expose or allow inference of the sources used in response generation were included. The analysis ran on prompts monitored in real production on the GEO Metrics platform, with a snapshot of the most recent available response per (query, engine) pair to avoid double counting.
The data is proprietary. The patterns that follow emerge from that sample.
Pattern 1: Three Citation Tiers — and One Engine That Never Cites
The volume of sources per response does not form a smooth gradient. It forms three clearly separated tiers.
At the top, alone: Grok. With an average of 45.6 sources per response (median 44, with peaks of up to 118 URLs in a single response), Grok has no competition in citation volume. It is a category of its own.
In the middle pack: AI Mode (13.8), ChatGPT (13.5), AI Overviews (10.3) and Perplexity (9.3). Four engines that cite regularly but with very different profiles from each other.
At the bottom: Copilot, with an average of 3.9 sources. It cites in 89% of its responses — almost always — but with very low volume per response.
And DeepSeek: 0.0. Zero sources, in 100% of its responses, across the full 14 months of analysis. Without variation. It is not that it cites rarely — it never cites at all.
Within the middle pack there are nuances that matter for strategy. Perplexity is the most consistent of all — standard deviation of just 2.2, meaning it almost always cites around 10 sources regardless of query type. ChatGPT is the most erratic — standard deviation of 11.6, with 22.7% of responses citing no source at all. For the same prompt, ChatGPT may respond with 0 sources or with 30.
Per-Engine Implications
Engine | Sources/resp | Consistency | Key note |
|---|---|---|---|
Grok | 45.6 | Medium | Up to 118 URLs in one response |
AI Mode | 13.8 | High | Social-first |
ChatGPT | 13.5 | Low | 22.7% with no source at all |
AI Overviews | 10.3 | High | Mirror of AI Mode |
Perplexity | 9.3 | Very high | σ = 2.2 — the most predictable |
Copilot | 3.9 | High | 89% of responses with source; +140% growth |
DeepSeek | 0.0 | Perfect | 0 sources across 14 months without exception |
Pattern 2: Each Engine Has a Different Source Diet
How much an engine cites is only half the story. The other half — the one that actually changes your strategy — is what it cites.
Each engine has a characteristic diet of source types. And they are incompatible diets.
Grok, AI Mode and AI Overviews are social-first. 80.9% of Grok's responses cite social networks. 66.5% include a YouTube video. The two Google engines (AI Mode and AI Overviews) replicate that social profile with slight variations — not by coincidence, they share a base index.
ChatGPT is the editorial engine. Wikipedia appears in 15.8% of its responses and news media in 13.6%. These are the highest editorial proportions in the study. ChatGPT is practically the only engine that consistently covers that ground.
Perplexity distributes. Balanced across video, community, professional networks and editorial sources. It is the engine with the most diversified diet — which also makes it the least predictable in terms of what type of content it prioritizes for each query.
Copilot goes niche. It barely touches mainstream sources. It draws from niche comparison sites through the Bing index. For Copilot, what matters is not Wikipedia or YouTube — it is appearing in the right comparison site.
The Direct Consequence
The YouTube and Instagram content that feeds Grok and AI Mode does not give you a single citation in ChatGPT. The Wikipedia and media presence that drives ChatGPT is almost irrelevant for Grok. There is no universally "good content" — there is good content per engine.
A GEO strategy that optimizes for a single source type is optimizing for a subset of engines — and ignoring the rest.
Pattern 3: Parallel Realities — They Share 1 in Every 10 Sources
If two engines drew from the same pool, their responses would at least be supported by the same evidence. That is not the case.
We calculated the Jaccard similarity between the sets of domains that each pair of engines cites for the same query. The average is ~10%: for the same question, two engines share on average 1 source domain out of every 10. The typical pair shares none.
The pairs with the highest similarity are those most architecturally related:
Engine pair | Jaccard | Why |
|---|---|---|
AI Mode × AI Overviews | 24.6% | Google family, shared index |
Grok × AI Mode | 16.8% | Similar social profile |
Grok × AI Overviews | 15.3% | Same social bias |
ChatGPT × Perplexity | 11.2% | Both with editorial component |
Perplexity × Copilot | 8.9% | No structural overlap |
ChatGPT × Copilot | 7.4% | Independent architectures |
ChatGPT × Grok | 6.1% | Opposite diets |
Even the pair with the highest overlap — AI Mode and AI Overviews, the Google family — shares only 24.6% of its sources for the same query. The rest share between 6% and 16%.
The implication is direct: two engines answering the same question are working with nearly completely different evidence. Their responses don't diverge by chance — they diverge because they are drawing information from different source universes.
Pattern 4: The Full Profile of Each Engine
Grok — Massive Social
45.6 sources per response · 80.9% of responses with social networks · 66.5% with YouTube
The engine with the highest citation volume by a wide margin. Its diet is social — X, YouTube, creator content. For a brand that wants citation in Grok, the strategy runs through active presence on social platforms with high volumes of indexable content.
ChatGPT — Editorial With Noise
13.5 sources per response · 15.8% Wikipedia · 13.6% news media · 22.7% with no source
The most erratic engine in the study. When it cites, it does so from high-authority editorial sources. When it doesn't cite, there is no visible signal as to why. Wikipedia remains the highest-return investment for ChatGPT — it is the only citation asset with cross-engine traction and the one with the most impact on editorial citation frequency.
AI Mode — Social and Structured
13.8 sources per response · 28.0% social networks · Google profile
The conversational version of the Google ecosystem. Social-first but more structured than Grok. Its source base has the highest overlap with AI Overviews — suggesting that a single source strategy can provide simultaneous coverage across both Google engines.
AI Overviews — Social and Brief
10.3 sources per response · 28.0% social · mirror of AI Mode
The other half of the Google family. Shares AI Mode's social profile with a lower source volume. The AI Mode/AI Overviews pair has the highest overlap in the entire study — 24.6% Jaccard — suggesting that an action that moves sources in one has a reasonable probability of moving them in the other.
Perplexity — Balanced and Predictable
9.3 sources per response · σ = 2.2 · diversified diet
The most consistent engine in the study. Almost always cites around 10 sources regardless of query type. Distributes across video, community, professional networks and editorial sources. Perplexity's predictability makes it the easiest engine to monitor and the most sensitive to short-term PR actions — because it crawls the web in real time with high consistency.
Copilot — Niche and Growing
3.9 sources per response · cites in 89% · +140% growth
The engine with the lowest citation volume among those that cite, but the fastest growing. Favors niche comparison portals through the Bing index. For Copilot, the strategy does not run through Wikipedia or YouTube — it runs through appearing in the right comparison site or niche directory. With high penetration in Microsoft corporate environments, it is the B2B engine par excellence.
DeepSeek — No Citations, Only Training
0.0 sources · 100% of responses without source · 14 months without variation
A case apart. DeepSeek never cites. Its brand visibility depends exclusively on training data — what appears in its corpus before the knowledge cutoff. There is no link equity that enters, no PR that impacts in real time. For DeepSeek, the only viable strategy is long-term entity building in the sources that form its training corpus.
The Conclusion That Changes Your GEO Strategy
There is no "the AI." There are seven engines that draw from nearly independent universes of evidence — and share barely 1 source out of every 10.
Three operational consequences:
1. Wikipedia is the only citation asset with cross-engine traction. It is the highest ROI investment across engines. ChatGPT cites it in 15.8% of its responses. AI Mode and AI Overviews include it frequently. Even Perplexity uses it as an editorial source. For any engine with an editorial citation component, Wikipedia remains the most powerful authority signal available.
2. The chatgpt.com referral in your analytics is already a visibility KPI, not noise. If ChatGPT starts citing you as a source — even in 13.5% of its responses — referral traffic from chatgpt.com will be visible in GA4. That number is not an analytics system anomaly — it is a direct signal of editorial citability.
3. Re-audit every quarter. Engines rewrite their citation rules in months, not years. Grok's source profile today will not be the same in six months. Copilot is growing 140% in citation volume. ChatGPT can add or remove sources erratically from one response to the next. A GEO strategy that is not monitored in real time is operating on data that is already outdated.
The question that defines GEO in 2026 is no longer just where does my brand rank in AI responses?. It is: in which engines am I cited as evidence, and what type of content drives that citation?
Analysis based on a representative sample of the GEO Metrics platform database. Proprietary data, prompts monitored in real production.
Monitor which sources AIs cite about you → trygeometrics.com
GEO & AEO expert focused on making brands visible inside AI-generated answers. He leads GEO Metrics, measuring how models like ChatGPT and Gemini cite, rank, and describe brands. His work helps companies move from SEO rankings to true visibility in AI-driven search.
See more articles
Learn actionable strategies, proven workflows, and expert tips to help your brand thrive.
Subscribe to GEO Metrics newsletter!
Receive expert advice, updates, and smart analytical insights directly in your inbox.










