ChatGPT search, Perplexity and Google’s AI Overviews each reach your page through a different door, and each door has its own key. ChatGPT search needs OpenAI’s search crawler, OAI-SearchBot, to be allowed. Perplexity needs its crawler, PerplexityBot. Google’s AI Overviews and AI Mode use the ordinary Google Search index, so they need Googlebot, a page that can be indexed, and permission to show a snippet. Opening one door does nothing for the other two.
A few words first, because the rest depends on them. An answer engine is a service that writes an answer to a question and names its sources, instead of only showing a list of links. A crawler is the program a company sends to fetch web pages. robots.txt is the text file at the root of your site that tells crawlers, by name, what they may fetch. An index is the stored copy of pages that a search engine looks things up in.
The three doors in one table
Everything in this table comes from the vendors’ own documentation, read on 24 September 2026.
| ChatGPT search | Perplexity | Google AI Overviews and AI Mode | |
|---|---|---|---|
| What feeds the answer | Pages crawled by OAI-SearchBot, plus other search providers OpenAI sometimes partners with | Pages crawled by PerplexityBot | Google’s own Search index |
| robots.txt name that controls it | OAI-SearchBot | PerplexityBot | Googlebot |
| Other switches you will meet | GPTBot (training), ChatGPT-User (fetches when a person asks) | Perplexity-User (fetches when a person asks) | Snippet rules, the Search Console generative AI setting, Google-Extended (not a Search switch) |
| What the docs say about links | Answers “may include citations”; a citation opens its source | PerplexityBot exists “to surface and link websites” | Eligible pages can be shown “as a supporting link” |
| How to check your access | Look for OAI-SearchBot in your server logs, from the IP ranges OpenAI publishes | Look for PerplexityBot in your logs, from Perplexity’s published IP ranges | URL Inspection in Search Console; the Generative AI performance report |
| How long a change takes | About 24 hours after a robots.txt change | Up to 24 hours | Several days to several months for snippet-control changes |
The last Google cell comes from its AI features page, which says recrawling “can take anywhere from several days to several months” after you change preview controls such as snippet rules.
Door one: ChatGPT search
OpenAI’s publishers FAQ starts from a generous premise: “Any public website can appear in ChatGPT search.” Its crawler page says OAI-SearchBot “is used to surface websites in search results in ChatGPT’s search features.” Sites that block it “will not be shown in ChatGPT search answers, though can still appear as navigational links.” OpenAI’s help page adds the second half of the key: allow the crawler and confirm that your host or content delivery network (CDN, the service many sites sit behind for speed and filtering) lets through OpenAI’s published searchbot IP addresses.
Two details matter. First, OpenAI’s help page says “ChatGPT search sometimes partners with other search providers”, rewriting the question into targeted searches, and it links Microsoft’s and Shopify’s privacy policies in that section. Your own crawl access is one input, not the only one. Second, GPTBot is a different door: it collects content that may be used to train OpenAI’s models. Blocking it does not take you out of ChatGPT search. GPTBot vs OAI-SearchBot vs ChatGPT-User goes through all three OpenAI names.
Door two: Perplexity
Perplexity’s crawler page says PerplexityBot “is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models.” It recommends allowing the bot in robots.txt and letting through requests from its published IP ranges, and it gives step-by-step firewall settings for Cloudflare and AWS.
Its second agent, Perplexity-User, visits a page when a person asks a question and may “include a link to the page in its response.” Perplexity says that because a person requested the fetch, it “generally ignores robots.txt rules.” So a robots.txt block on Perplexity-User is not a reliable way to keep a page out. If a page must stay private, put it behind a login.
Door three: Google AI Overviews and AI Mode
Google has no separate AI crawler for these features. Its AI features page says “robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search.” To be eligible as a supporting link, “a page must be indexed and eligible to be shown in Google Search with a snippet”, and Google adds that “There are no additional technical requirements.”
Three extra switches sit around that door. The snippet rules nosnippet, max-snippet and data-nosnippet, plus the indexing rule noindex, limit what Google may show or use. Search Console has a generative AI setting where include is the default; Google says it had reached all websites by 31 August 2026. And Google-Extended, a robots.txt token for Gemini training and grounding, “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”
Even with every switch right, you may see nothing. Google says AI Overviews “often don’t trigger”, because they only appear when its systems judge them useful on top of the ordinary results.
What does no vendor tell you?
How it chooses. OpenAI’s help page is the most direct: “ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information. Placement is not guaranteed.” Google says its AI features are rooted in its core Search ranking systems and that indexing and serving are not guaranteed. Perplexity’s crawler page says nothing about selection at all.
So when someone offers you “the ChatGPT citation formula”, they are selling a guess. What you can control is covered in How answer engines choose sources, and the page work in How to get cited by ChatGPT. Opening the doors is the part of AEO (answer engine optimization) that you can check with certainty.
Four myths and the facts
| Myth | Fact, from the vendor’s own page |
|---|---|
| ”Blocking GPTBot keeps me out of ChatGPT.” | GPTBot is the training crawler. ChatGPT search uses OAI-SearchBot. |
| ”Blocking Google-Extended keeps me out of AI Overviews.” | Google says Google-Extended does not affect inclusion in Google Search. AI Overviews follow Googlebot, snippet rules and the Search Console setting. |
| ”A robots.txt line for Perplexity-User keeps Perplexity away.” | Perplexity says Perplexity-User generally ignores robots.txt, because a person asked. |
| ”If I let all the bots in, I will be cited.” | Access makes you eligible. OpenAI: “Placement is not guaranteed.” |
A worked example (hypothetical)
Imagine a Dutch company that sells accounting software. Its hosting company switches on stricter bot protection one spring. Google traffic carries on as normal, and the page on invoicing rules still shows up in AI Overviews, because Googlebot is let through.
But the server logs show no OAI-SearchBot or PerplexityBot visits for weeks, only blocked requests. robots.txt allows both bots. The firewall does not. The fix is a rule that lets through the IP ranges OpenAI and Perplexity publish. Nothing about that fix promises a citation. It only puts two of the three doors back on their hinges.
The 15-minute access check
- Minutes 0 to 3. Open
https://your-site/robots.txt. Find the groups that name OAI-SearchBot, PerplexityBot and Googlebot, and theUser-agent: *group. Write down open or closed for each. - Minutes 3 to 6. Open your firewall or CDN settings. Look for bot protection, challenge pages or country blocks that could stop the crawlers before they read robots.txt.
- Minutes 6 to 10. Search the last seven days of server logs for
OAI-SearchBot,PerplexityBotandGooglebot. Note the status codes: 200 means served, 403 means refused. Check a few addresses against the IP lists OpenAI and Perplexity publish, because anyone can fake a user-agent name. - Minutes 10 to 13. In Search Console, run URL Inspection on your most important page. Then open the generative AI setting and confirm it says include.
- Minutes 13 to 15. Write three lines: ChatGPT door, Perplexity door, Google door. Each one is open, closed or unknown.
If a line says unknown, that is your next job. The book that treats this as one system, The Findability Engine, is in production and not for sale yet. The check above needs no book: a text file, a firewall screen and a log.
Cite this:ChatGPT, Perplexity, Google AI Overviews: three engines, three doors to your page.Len P. van der Hof. https://lenvanderhof.com/en/blog/three-answer-engines-three-doors/ ·