← All articles

Being cited by an AI happens 99% somewhere other than your site

ChatGPT names a brand in only a tiny fraction of its answers, and the sources it favours do not belong to you. What that changes for a company that wants to exist inside those answers.

Dots Papers cover for the article on being cited by AI assistants and generative engines

In short

  • Assistants very rarely name brands. Across 34,234 answers recorded between 14 January and 13 February 2026, ChatGPT displayed a citation in 0.59% of answers, Perplexity in 13.05% and Grok in 27.01%. The gap runs from 1 to 22 between ChatGPT and Perplexity, and from 1 to 46 between ChatGPT and Grok.
  • The dominant sources are not company websites. Reddit is the leading source on Perplexity and Google AI Overviews; on ChatGPT it is Wikipedia, which accounts for 47.9% of citations among the top ten sources. Journalism accounts for 27% of cited sources.
  • One platform does not predict another. Only 11% of domains cited by ChatGPT are also cited by Perplexity: these are two different logics, not one shared ranking.
  • And it all moves fast. Reddit’s share of ChatGPT citations fell from roughly 60% to 10% in six weeks in late 2025, after a single upstream configuration change. It happened again in August 2026: from 3.8% to 0.5% in a week, after a change in how ChatGPT Search queries the web.

The question almost always arrives in this form: what do we need to change on our website so ChatGPT cites us? It is a fair question, and it contains a false assumption. Most of what decides your presence in a generated answer is not on your website.

That is not a reason to do nothing. It is a reason to do something other than what is usually sold to you.

Update, 24 August 2026 On 8 August, ChatGPT Search changed how it queries the web: its internal queries targeting a named site, using the “site:” operator, jumped from 0.4% to nearly 17% in a single day. As a result, Reddit’s share of its citations collapsed from 3.8% to 0.5% by 14 August, a relative drop of 86%. The move is specific to ChatGPT, as the decline on Google AI Overviews and AI Mode remains gradual and far smaller, and OpenAI denies any deliberate targeting. The lesson is twofold: the volatility described in this article is anything but theoretical, and scoped queries now favour official documentation and help centres, ground that does belong to you. In the same month, ChatGPT opened advertiser sign-up across Europe, a second regime alongside citations: the urgent part there is not the buying but the measurement.

What the data says, and why it surprises

The reference analysis on this subject covers 680 million citations recorded between August 2024 and June 2025 across ChatGPT, Google AI Overviews and Perplexity. Three findings come out of it, all counter-intuitive.

User-generated content dominates. Reddit is the leading source on Perplexity, where it accounts for 46.7% of citations among the top ten sources, and on Google AI Overviews, at 21.0%. On ChatGPT, the leading source is Wikipedia, at 47.9%. The common thread: these are places where people describe an experience, not pages where a company describes its offer.

The press matters, but intermittently. Journalism accounts for 27% of citations, rising to 49% on time-sensitive questions. In other words, the more a question touches the news, the more the citation goes to a media outlet.

Brands are rarely named. This is the most useful figure for arbitrating a budget: where Perplexity displays a citation in 13.05% of its answers, ChatGPT does so in 0.59% of cases. One caveat applies, and it points the same way: this measure counts displayed links, not brand mentions. An assistant can name you without citing you, and nobody measures that case. Building a strategy on the opposite assumption means investing in a channel that, on the most used platform, almost never opens.

The mistake that costs the most Buying a service that promises to get you cited by ChatGPT, measured on a sample of questions chosen by the provider. On a platform that names a brand in under one per cent of answers, manufacturing a positive demonstration is trivial and drawing a conclusion from it is impossible. Ask how many questions the measurement covered, and who chose them.

One platform is not like another

This is the point that invalidates most dashboards sold on this subject: only 11% of domains cited by ChatGPT are also cited by Perplexity. There is no single ranking whose rungs you climb.

Platform What it favours What that means for you
ChatGPT Community sources and major media, brands rarely named Presence comes through third parties, not through your pages
Perplexity Cites far more widely, company websites included This is where the quality of your own pages shows up most
Google AI Overviews Strong video bias, YouTube far ahead A filmed answer can be worth more than a well-written page
All of them High volatility, over weeks Measuring once tells you nothing, you need a series
Reading of data published in 2026. The orders of magnitude matter more than exact values, which move quickly.

On volatility, two examples are enough. In late 2025, Reddit’s share of ChatGPT citations fell from roughly 60% to 10% in six weeks, following a configuration change at an upstream supplier. In August 2026, it collapsed from 3.8% to 0.5% in a week, when ChatGPT Search started querying named sites directly rather than the open web. In both cases, nobody at Reddit had changed anything about the content.

What remains under your control

Three things, in decreasing order of return.

  1. Exist where people talk about you. Reviews, community discussions, comparisons, video answers. It is uncomfortable because it is not your ground, and that is where most of it is decided. This presence is not bought, it is earned by being publicly useful.
  2. Make your pages extractable. A direct answer at the top, dated data, an identifiable entity, clean markup. It does not get you cited by magic, but it makes you citable when the question lands on your ground, which is the case for precise, commercial questions. The August 2026 shift strengthens this lever: ChatGPT’s scoped queries primarily target official documentation and help centres, in other words pages you own.
  3. Allow access. Some AI crawlers are blocked by default by security rules or by a web application firewall, without anyone having decided it. It is the least glamorous check and the most profitable: you cannot be cited by an engine that cannot read the page.

Twenty-minute self-check: list the ten questions a customer actually asks you before buying, put them to ChatGPT and Perplexity, and note who gets cited. Repeat three weeks later. It is not a scientific measurement, but it is infinitely more honest than a report built on questions chosen to please you.

The options, and what they are worth

Option What you get Who it suits
Check AI crawler access Near-zero cost, binary effect. Either you are readable or you are not Everyone, and it is the first move
Structure pages for extraction Improves citability where your site stands a chance: precise questions, niche subjects Sites with real substantive content
Work on presence with third parties The strongest lever according to the data, and the slowest. Nothing is bought, everything is built Brands willing to talk publicly about their trade
Buy a guaranteed-citation service The guarantee cannot be kept: neither you nor the provider controls a model’s ranking No case at all
The first three stack. The fourth is the most frequent source of disappointment on this subject.

Tooling

Two needs can be handled by tooling. Knowing who actually reaches your pages, AI crawlers included, and checking that your content is structured as an answer rather than as a brochure.

On PrestaShop we deploy the modules published by Datafirefly Limited, our agency’s sister company: AI Crawler Manager to see and arbitrate AI crawler access, and Semantic Audit to check page structure. One-off purchase · 12 months of updates.

Method beats tooling, and the imbalance is particularly clear here: no module makes you exist inside a community discussion, and that is where most citations come from.

What this says more broadly

Classic search rested on an implicit promise: doing your job well on your own website was enough to be found. That is no longer true, and it is not a 2026 novelty: reviews, marketplaces and social platforms had already moved part of visibility off your own ground.

What has changed is the scale. A company that only talks about itself on its own website accepts being absent from the layer where part of the decisions are now taken. It is the heart of our artificial intelligence practice, and it is a question of editorial strategy long before it is a technical one. We give the pragmatic version in our article on what actually works with AI in a small company.

Sources

The figures in this article point back to the study that produced them, accessed on 11 August 2026. We do not cite a source we have not read.

  1. Superlines, AI Search Statistics 2026, March 2026, 34,234 answers recorded across 10 platforms from 14 January to 13 February 2026. Read
  2. Profound, AI Platform Citation Patterns, June 2025, 680 million citations across ChatGPT, Google AI Overviews and Perplexity, August 2024 to June 2025. Read
  3. Profound, Answer Engine Citation Overlap Strategy, July 2025, 100,000 identical prompts put to ChatGPT and Perplexity. Read
  4. Semrush, The Most-Cited Domains in AI, November 2025, 230,000 prompts and more than 100 million citations, 14 July to 12 October 2025. Read
  5. Muck Rack, Generative Pulse, What Is AI Reading?, May 2026, more than 25 million links from ChatGPT, Claude and Gemini across 17 industries. Read
  6. Promptwatch, Why Did ChatGPT Stop Citing Reddit, August 2026, daily share of ChatGPT Search citations from 7 July to 17 August 2026. Read

FAQ

Do I need an llms.txt file?

It costs nothing and clarifies what you allow, but no public evidence shows it triggers citations. Treat it as a technical courtesy, not a lever. The equivalent gesture that does have a measurable effect is checking that your pages are reachable by crawlers.

My site is never cited, is that a technical problem?

Rarely. On a platform that names a brand in under one per cent of answers, absence is the normal case. Check crawler access first, then look at whether your competitors are cited: if they are not either, there is nothing to fix.

How do I measure this seriously?

As a series, never as a snapshot. Ten to twenty questions your customers really ask, recorded at regular intervals, on at least two platforms. A single reading tells you nothing given the volatility observed.

Should I block AI crawlers?

It is a trade-off, not an obvious call. Blocking protects your content and removes you from answers. Allowing makes you citable and feeds models. The right answer depends on your business model, and it deserves to be taken knowingly rather than inherited from a firewall setting.

Can Dotsland help?

Yes. A citability reading on your real customer questions across several platforms, a crawler access check, page structuring for extraction, and arbitrage on third-party presence. It is the heart of our artificial intelligence practice. Let’s talk.

Want to apply this to your own business?

Get in touch →

Further reading

Leave a comment

Your email address will not be published. Required fields are marked *

16 − 15 =