Technical GEO is the plumbing under generative engine optimization. GEO is the work of making your pages usable as sources in AI-written answers, such as ChatGPT search, Perplexity or Google’s AI Overviews. It has two halves. The writing half decides whether a page is worth quoting. The technical half decides whether a machine can reach and use the page at all: crawler access, HTML it can read, indexing, snippet permissions, and a stable address with a date and an author.
A developer fixes the second half, and it comes first. The best paragraph on a page is worth nothing to a system that never receives it.
What does GEO stand for, and what makes it technical?
GEO stands for generative engine optimization. Google’s own guide to its AI search features uses the term, then adds a line worth pinning above your desk: “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”
That is the myth and the fact in one sentence. The myth: AI answers need a new technical stack. The fact: the doors are mostly the old doors, with a few new names on them. “Technical” means the parts that live in server settings, page templates and HTTP headers (the hidden lines a server sends along with every page), not in your sentences. For the wider definition, see GEO meaning.
Why does the plumbing come before the writing?
Because each engine documents a door, and a closed door ends the conversation.
- Google: to appear as a supporting link in AI Overviews or AI Mode, “a page must be indexed and eligible to be shown in Google Search with a snippet”. Google adds: “There are no additional technical requirements.”
- ChatGPT search: OpenAI’s help page says to allow OAI-SearchBot and to “confirm that the website host or content delivery network allows traffic from OpenAI’s published searchbot IP addresses.”
- Perplexity: its crawler page recommends allowing PerplexityBot in robots.txt and “permitting requests from our published IP ranges”.
None of these pages promises a citation once the door is open. They only tell you how to avoid being ruled out.
The seven-point checklist
This is the table to screenshot. Each row has a check you can run in about two minutes.
| # | Check | Why it matters | Two-minute check |
|---|---|---|---|
| 1 | Each AI search crawler may fetch the page | A blocked crawler never sees the text | Read your live robots.txt; check your firewall or CDN bot settings |
| 2 | The words are in the HTML the server sends | Not every bot runs JavaScript | Fetch the raw page and search it for one key sentence |
| 3 | The page is indexable, returns 200, and has one canonical address | Google’s AI features draw on its Search index | Check the status code, look for noindex, read the canonical tag |
| 4 | Snippet rules allow quoting | nosnippet also blocks use in AI Overviews and AI Mode | Search the HTML and headers for nosnippet and max-snippet |
| 5 | Stable address, visible date, named author | Moved pages break old links; readers check who and when | Open the page: is there a date and a byline? Do old addresses redirect? |
| 6 | Structured data matches the visible page | It helps machines label things; it is not a key | Compare the markup’s author and dates with what the page shows |
| 7 | llms.txt, if any, is a deliberate choice | Google Search ignores it | Decide whether a named tool reads it; otherwise skip it |
1. Crawler access, one bot at a time
A crawler is a program a company sends to fetch web pages. robots.txt is a small text file at the root of your site that tells crawlers, by name, which paths they may fetch. The companies below run separate crawlers for separate jobs, so “allow AI” is not one setting:
| Company | Search crawler (answers and links) | Training crawler | Fetches when a person asks |
|---|---|---|---|
| Googlebot, which also serves AI Overviews and AI Mode | Google-Extended (a robots.txt token for Gemini training and grounding outside Search, not a separate crawler) | Not covered here | |
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User |
| Perplexity | PerplexityBot | None listed: Perplexity says PerplexityBot is not used for training foundation models | Perplexity-User |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User |
Two consequences. Blocking a training crawler does not remove you from search answers, and Google says Google-Extended “does not impact a site’s inclusion in Google Search”. And the person-triggered fetchers play by different rules: OpenAI says robots.txt “may not apply” to ChatGPT-User, and Perplexity says Perplexity-User “generally ignores robots.txt rules.” GPTBot vs OAI-SearchBot vs ChatGPT-User walks through the OpenAI side in detail.
Then look past the file. Google’s AI features page lists this among its basics: “Ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure”. A CDN (content delivery network, the service that sits in front of many sites to speed them up and filter traffic) can block a bot that robots.txt allows. OpenAI says its search systems need about 24 hours to adjust after a robots.txt change. Perplexity says up to 24 hours.
2. The words must be in the HTML
Some sites send an almost empty page and build the content in the visitor’s browser with JavaScript (the programming language that runs inside web pages). Google can process that, but its own guide says server-side or pre-rendering “is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript.”
The two-minute check uses curl, a free command-line tool that fetches a page the way a simple crawler does:
curl -s https://example.com/your-page | grep -c "a sentence you expect on the page"
A result of 0 means the sentence is not in what the server sent. Fix that before anything else on this list.
3. Indexable, reachable, one address
Google’s minimum for Search is short: Googlebot is not blocked, the page returns HTTP 200 (the code for “this worked”), and it has indexable content. A noindex tag takes the page out of Search entirely, and with it out of AI Overviews and AI Mode.
The canonical tag (rel="canonical") names the one preferred address when several addresses show the same content. Google recommends that the canonical page names itself. Check the status with curl -sI, then search the HTML for noindex and canonical. Google’s URL Inspection tool in Search Console shows what Googlebot received.
4. Snippet rules decide what may be quoted
A snippet is the short piece of text a search result may show. Google’s specification says nosnippet “will also prevent the content from being used as a direct input for AI Overviews and AI Mode”, and max-snippet limits how much may be used. data-nosnippet does the same for one part of a page. These rules are useful on purpose and harmful by accident, for example when an old template sets them on every page.
Google adds one more switch. Search Console now has a Search generative AI setting (Settings > Search generative AI), rolled out to all websites as of 31 August 2026. Include is the default. If someone chose exclude, your site’s links will not appear in those features.
5. Stable address, visible date, named author
If a page moves, answers that linked the old address send readers to an error page. Google recommends a permanent server-side redirect, calling it “the best way to ensure that Google Search and people are directed to the correct page.”
Google also asks for a visible date: “Add a user-visible date to the page and feature it prominently.” And for authorship: “We strongly encourage adding accurate authorship information, such as bylines to content where readers might expect it.” No engine says a date or a byline earns a citation. They are on this list because a person checking an AI answer uses them to decide whether to trust the page they land on.
6. Structured data helps, but it is not a key
Structured data is code in the page, usually schema.org markup, that labels things like the author and the publish date for machines. Google is blunt about its AI features: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add.” It still recommends structured data for rich results, and asks that it match the visible text. Markup that claims an author the page does not show is a liability, not a shortcut.
7. llms.txt is optional
llms.txt is a proposed file that lists a site’s key pages for AI tools. Google’s guide says Google Search ignores such files, and that creating one “will neither harm nor help your site’s visibility or rankings in Google Search”. Neither OpenAI’s nor Perplexity’s crawler page lists it among the controls for their crawlers. What is llms.txt? covers when it is still worth keeping.
A worked example (hypothetical)
Imagine a small insurance broker in Rotterdam with a page comparing boat-insurance premiums. The page looks fine in a browser. But the premium table loads with JavaScript after a cookie banner, and curl returns only a heading and the word “Loading”.
A crawler that reads only the raw HTML gets nothing from that page to use when someone asks an AI assistant about boat-insurance costs. The broker’s developer renders the table on the server, so the numbers are in the HTML. Nothing about that fix guarantees a citation. It only puts the page back in the running, which is the whole job of technical GEO.
What technical GEO will not do
It will not buy a ranking. OpenAI’s help page says ChatGPT “ranks search results using multiple factors” and that “Placement is not guaranteed.” Google says indexing and serving are not guaranteed either. No vendor publishes a formula you can tune.
It will not make a page worth quoting. That is the other half of GEO: a clear answer near the top, evidence a reader can check, something the ten other pages on the topic do not say. How answer engines choose sources covers that half. The book that treats both halves as one system, The Findability Engine, is in production and not for sale yet. The GEO glossary entry has the short definition.
Try this today (20 minutes)
Pick the one page you most want an AI answer to use.
- Minutes 0 to 5. Open
https://your-site/robots.txt. Find the groups for Googlebot, OAI-SearchBot and PerplexityBot, and theUser-agent: *group. Write down whether each may fetch your page’s path. - Minutes 5 to 10. Run the
curl | grepcheck above with one sentence from the page. Then runcurl -sIand note the status code. - Minutes 10 to 15. Search the page source for
noindex,nosnippet,max-snippetandcanonical. Note anything you did not put there on purpose. - Minutes 15 to 20. Check that the page shows a date and an author. In Search Console, open Settings > Search generative AI and confirm the setting says include.
Any “no” on this list is a fix for a developer, not a brief for a copywriter. Fix the valve before you polish the tap.
Cite this:Technical GEO: the plumbing that decides whether AI can use your page.Len P. van der Hof. https://lenvanderhof.com/en/blog/technical-geo/ ·