01Two questions that get confused
Being named and being fetched are separate. An assistant can name your app from what it absorbed during training, without ever requesting your site. It can also fetch your site at answer time and still not name you. Measuring one tells you nothing about the other, and most advice on this subject silently merges them.
The first question we already measure and publish, including the zeroes. This page is the second one: over a fixed window, which crawlers actually requested pages from this domain.
02What arrived, 14 days to 2026-09-21
| User agent | Operator | Role | Requests |
|---|---|---|---|
| Amazonbot | Amazon | search index | 1140 |
| ClaudeBot | Anthropic | training | 901 |
| Googlebot | search index | 797 | |
| bingbot | Microsoft | search index | 435 |
| Applebot | Apple | search index | 121 |
| GPTBot | OpenAI | training | 67 |
| PerplexityBot | Perplexity | answer-time | 10 |
| Bytespider | ByteDance | training | 5 |
| OAI-SearchBot | OpenAI | answer-time | never |
| ChatGPT-User | OpenAI | answer-time | never |
| Claude-User | Anthropic | answer-time | never |
| CCBot | Common Crawl | training | never |
| meta-externalagent | Meta | training | never |
| Google-Extended | training | never | |
| Perplexity-User | Perplexity | answer-time | never |
| Claude-SearchBot | Anthropic | answer-time | never |
| DuckAssistBot | DuckDuckGo | answer-time | never |
| MistralAI-User | Mistral | answer-time | never |
03OpenAI's answer-time agents never arrived. Perplexity's did.
Operators publish separate agents for separate jobs: one that gathers pages for training, another that fetches a page while answering a live question. Over this window OAI-SearchBot, ChatGPT-User, Claude-User, Perplexity-User, Claude-SearchBot, DuckAssistBot, MistralAI-User registered zero requests — both of OpenAI's answer-time agents among them. PerplexityBot did arrive, 10 times.
So the honest version is not "assistants never fetch" — one of them does, at a volume of about one request a day. It is narrower than that: for this domain, in this window, no agent identifying itself as OpenAI's fetched this site while answering. We first read that as "ChatGPT cannot see our pages". The next section is why that reading was wrong.
The obvious objection is that we blocked them. We did not: OAI-SearchBot, ChatGPT-User, Claude-User, Perplexity-User, Claude-SearchBot, DuckAssistBot, MistralAI-User are each named in our robots.txt with Allow: /, alongside the training crawlers. The zero is their behaviour, not our configuration — check your own robots file before reading your own zero as a finding.
04Zero fetches, and ChatGPT linked to us anyway
On 21 September 2026 we asked ChatGPT, in a temporary chat with memory off, "free tool to check my iOS app's keyword ranking in the App Store". It listed our free rank checker second of five and linked straight to /tools/rank-check, with utm_source=chatgpt.com appended. Our own visit log had been recording arrivals from chatgpt.com for two weeks — two of them went on to create an account.
So the access-log zero and the citation are both true at once. The answer was assembled from a search index ChatGPT already held, not from a live request by an agent announcing itself as OpenAI's. For a site owner the lesson is practical: a zero in your log does not mean you are invisible to ChatGPT, and a citation does not mean it read your page today.
Which index? OpenAI's help centre says ChatGPT search "sometimes partners with other search providers" and points to Microsoft's privacy statement for them — that is Bing. bingbot requested pages here 435 times over the window, all served normally — yet the same day, a Bing search for the exact title of our free-tools comparison returned nothing from this domain. We cannot tell which index produced the citation; we can tell that crawling and indexing are separate steps, and that the second one is where a new site is usually missing.
05A third of all crawler traffic goes to the plain-text twins
Every page here also exists as a Markdown twin at .md.txt, alongside llms.txt and llms-full.txt. Over the same window crawlers requested those twins 788 times out of 2244 AI requests in total — 35% of everything they asked for. HTML content pages took 1083.
We built that layer on the assumption it would be used and had no evidence either way. It is used heavily — not more than the HTML, but at the same order of magnitude, from a layer that costs almost nothing to publish. On this evidence it is not decorative.
06What this does and does not suggest doing
Worth doing
Publish a clean plain-text copy of your pages, keep your own description factual, and make sure the pages that describe your app are the ones a training crawl would find — third-party write-ups more than your own marketing. Then measure again in a month rather than trusting the change.
Get into the search indexes the assistants borrow. Verify the site in Bing Webmaster Tools as well as Google Search Console and submit the sitemap to both — crawled is not the same as indexed, and a page that is not in the index cannot be retrieved for an answer. Then search your own exact page titles there; it is the cheapest check there is.
Not worth doing
Rewriting pages weekly to chase an answer-time fetch that is not happening. Buying a tool that promises to "optimise for ChatGPT" without telling you whether the engine it measures ever requests your domain. And treating a single reading — ours or yours — as a trend.
Frequently asked
Does ChatGPT visit my website when someone asks about my app?
Sometimes, and you can check rather than assume. OpenAI documents separate agents for training and for answer-time browsing; the answer-time ones appear in your access log by name when they arrive. Over our own measured window they never did — and ChatGPT still linked to our pages, because its answer drew on a search index it already had rather than a live visit.
Is blocking AI crawlers in robots.txt a good idea?
If you want to be named in assistant answers, blocking the training crawlers works against that directly — they are the route by which your product becomes something the model knows. Blocking is a reasonable choice for other reasons, but it is a trade, not a free precaution.
Do llms.txt and Markdown copies of pages actually get read?
On this domain, heavily: 788 of 2244 AI crawler requests went to the plain-text twins over the measured window, against 1083 for the HTML pages. That is one site's evidence, not a rule, but the cost of publishing them is close to zero.
Can I fake my way into AI answers?
Not usefully. An assistant's answer is built from what has been written about you across many sources; a page asserting that you are the best does not become that. What moves it is being described accurately, in places other than your own domain, over time.
See this for your own app
Paste your store link, add your keywords. Rank, the apps above you and demand are measured every day.
Free plan: 1 app, 25 keywords in each of 2 storefronts, 30 days of history, no card, no end date. Pro is $20/mo.
Start free →