SEO, GEO and AEO Research guide

How Search and Answer Engines Choose Sources

Citation visibility is not won with an AI hack. It is earned by combining technical eligibility, clear answers, original value and evidence that survives scrutiny.

Five connected layers showing discoverability, interpretation, extraction, verification and citation value
A source must pass several gates before it can become useful evidence in an answer.

The hard truth: GEO is not a shortcut around SEO

The fastest way to waste a GEO budget is to treat generative search as a separate universe with secret rules.

It is not.

For Google, this is now explicit. Its June 2026 generative-search guidance says that AI Overviews and AI Mode are rooted in Google’s existing Search index, ranking systems and quality systems. A page must first be indexed and eligible to appear with a snippet. There is no special AI schema that bypasses those requirements.

OpenAI and Microsoft expose different controls and reporting, but the same practical sequence still applies:

  1. The system must be able to discover and fetch the page.
  2. It must understand what the page and its entities are about.
  3. It must find a passage relevant to the question.
  4. It must have enough reason to trust and attribute that passage.
  5. The passage must add more value than the alternatives.

That sequence is the real relationship between SEO, AEO and GEO.

  • SEO creates technical and editorial eligibility.
  • AEO makes answers clear enough to retrieve and present.
  • GEO focuses on whether generative systems select, use and cite the source.

The labels differ. The underlying work overlaps heavily.

Verified principle: optimize for a useful, crawlable and evidential page. Do not create an inferior page for an imagined robot reader.

The Citation Readiness Model

I use a five-layer model to turn a vague ambition such as “rank in AI” into work that can actually be audited.

LayerThe engine’s questionYour jobTypical failure
1. DiscoverableCan I access and index this URL?Remove technical barriers and expose stable URLsBlocked crawling, noindex, broken rendering or duplicate URLs
2. InterpretableWhat is this page, who created it and what does it concern?Make topics, entities, authorship and relationships explicitAmbiguous titles, inconsistent entities or misleading schema
3. ExtractableIs there a passage that directly resolves the question?Write clear answers inside a coherent articleThe answer is buried, vague or dependent on missing context
4. VerifiableCan the claims be checked?Add primary evidence, dates, methods and limitationsUnsupported statistics, anonymous claims or stale facts
5. Worth citingDoes this source contribute something useful and distinct?Publish original experience, analysis, data or synthesisCommodity content that repeats the existing consensus

This model is an editorial heuristic, not a disclosed ranking factor. Its value is operational: when visibility is weak, it tells you where to investigate instead of reaching for random “AI optimization” tactics.

Layer 1: become technically eligible before polishing prose

A brilliant answer on an inaccessible URL is invisible.

Check the response, not just the browser

For every important article, verify:

  • The canonical URL returns a stable 200 response.
  • The complete primary content is present in the rendered HTML without requiring a click, swipe, search or login.
  • Important CSS, JavaScript and images are not accidentally blocked.
  • The page is not carrying a noindex directive.
  • The canonical points to the intended URL.
  • Internal links use ordinary crawlable <a href> links.
  • The URL appears in the XML sitemap with an accurate modification date after substantive changes.

Google can render JavaScript, but its own guidance still notes that JavaScript SEO is more complex. Server-rendered or statically generated primary content removes a class of avoidable failure.

Do not confuse crawling with indexing

robots.txt manages crawler access. It is not a reliable removal mechanism.

If Google cannot crawl a blocked page, it cannot see a noindex directive on that page. The URL may still be known through links and may appear without a useful description. For confidential material, use authentication. For index removal, allow the crawler to see noindex or return the appropriate status.

That distinction sounds basic. It is still one of the most expensive technical SEO mistakes.

Use the right crawler control for the right objective

ControlPrimary purposePractical decision
GooglebotGoogle Search, including Search AI featuresAllow important public pages if you want Google visibility
Google-ExtendedControls use for specified Gemini training and grounding systems outside Google SearchDecide separately from Google Search access
OAI-SearchBotInclusion in ChatGPT search answersAllow if ChatGPT search visibility is desired
GPTBotPotential training of OpenAI foundation modelsSet according to your training-data policy; it is independent of Search
ChatGPT-UserUser-triggered visits and actionsDo not treat it as the ChatGPT Search indexing control
IndexNowNotifies participating engines that URLs changedTrigger on meaningful additions, updates and deletions; it is not a ranking boost

OpenAI’s crawler documentation is unusually clear: a publisher can allow OAI-SearchBot while disallowing GPTBot. Search participation does not require a blanket decision about model training.

The same separation matters at Google. Googlebot governs Search access. Google-Extended is a separate product token for specified Gemini uses and does not replace Googlebot.

Layer 2: make the page interpretable

Machines do not need decorative metadata. They need consistent facts.

Establish one canonical identity for each thing

An article should make these entities unambiguous:

  • The article itself
  • The author
  • The publishing website or organization
  • The primary topic
  • Products, organizations, people or places discussed
  • The language version and its translated counterpart

Use the same author name, profile URL and relevant sameAs references across the site. Link the article’s author to a substantive profile page. Keep role descriptions and biographies factual rather than inflating them with vague authority language.

For bilingual publishing, give each language its own URL and connect equivalents with reciprocal hreflang. Do not canonicalize the Dutch article to the English article: they are localized alternatives, not duplicates to collapse.

Use structured data as corroboration

Accurate BlogPosting or Article, Person, ProfilePage, WebPage and BreadcrumbList markup can help systems interpret the page. Include supported properties such as:

  • headline
  • description
  • datePublished
  • dateModified
  • author
  • author.url
  • Representative images
  • The canonical article URL

But the markup must agree with what readers see. Do not invent credentials in JSON-LD, hide marked-up content or add schema for content that is absent from the page.

Structured data is a clarification layer, not a citation switch. Google explicitly says correct markup does not guarantee a rich result, and no special structured data is required for its generative Search features.

Make updates credible

An “updated” date should mean something changed materially.

When revising a high-stakes guide:

  1. Update inaccurate claims.
  2. Recheck outbound evidence.
  3. Add a visible modification date.
  4. Record a short change summary.
  5. Update sitemap lastmod.

Changing the date without changing the substance creates freshness theatre, not trust.

Layer 3: create answer assets, not answer fragments

A direct answer is useful. Artificially chopping every paragraph into tiny “AI chunks” is not.

Google now explicitly says that forced chunking is unnecessary. Its systems can understand nuanced pages, and there is no ideal word count. The right unit is the amount of explanation a reader needs.

The practical compromise is an answer asset: a passage that can stand on its own while remaining part of a deeper argument.

The four-part answer asset

For an important question, write:

  1. Answer: state the conclusion in one or two sentences.
  2. Conditions: explain when the answer applies and when it does not.
  3. Evidence: name the source, method, standard or observation.
  4. Next action: tell the reader what to do with the information.

Example:

Does Article schema improve AI citations? It can clarify the page’s type, author and publication details, but it does not guarantee citation. Use valid markup that matches visible content, validate it, and judge success through indexing and visibility data rather than assuming the schema caused selection.

That paragraph is concise enough to quote and complete enough not to mislead.

Match headings to genuine subproblems

Google says its AI features may use query fan-out: several related searches are issued to explore the subtopics behind a complex question.

The wrong response is producing dozens of thin pages for every imagined prompt variation. Google specifically warns against scaled content created to manipulate rankings.

The better response is to cover the natural decision path inside one coherent resource:

  • What is the concept?
  • Who is it for?
  • What must be true first?
  • How is it implemented?
  • What can go wrong?
  • How is it measured?
  • What are the limitations?

Use descriptive headings for those real information needs. Then create a separate page only when a subtopic deserves a distinct search intent, substantial depth and its own internal-link destination.

Prefer explicit language over clever ambiguity

Write “OAI-SearchBot controls ChatGPT Search crawling” before using “the bot.” Write “the page was updated on June 14, 2026” rather than “recently updated.” Define acronyms once.

Clear prose helps readers, translators, search systems and answer engines at the same time.

Layer 4: make claims verifiable

Generative systems can produce fluent text without reliable provenance. Your competitive advantage is not more fluency. It is auditability.

Build every important passage around Claim, Evidence and Entity

Use this editorial test:

ElementQuestionStrong implementation
ClaimWhat exactly are we asserting?A bounded statement with conditions, date and scope
EvidenceHow can a reader check it?A primary source, reproducible method or clearly labeled first-hand observation
EntityWho or what does the claim concern?An explicit product, organization, standard, person or dataset

“Schema helps AI” fails all three columns.

“Google states that structured data is not required for its generative Search features, although valid markup can still support rich-result eligibility” is bounded, attributable and tied to a specific platform.

Use a source hierarchy

For technical recommendations, prefer:

  1. Current first-party platform documentation
  2. Standards and specifications
  3. Peer-reviewed or directly inspectable research
  4. Reproducible first-party tests and data
  5. Expert analysis
  6. Unverified commentary

Lower-tier sources can identify questions. They should not silently become evidence for high-confidence claims.

Treat research results as research results

The original GEO paper found that citations, quotations and statistics improved source visibility in its benchmark, with gains of up to 40% in some experimental settings. That is valuable evidence, but not a universal instruction to stuff every paragraph with numbers.

The defensible conclusion is narrower:

  • Attributable evidence can improve usefulness and citation potential.
  • Effects differ by domain and engine.
  • A benchmark result is not a disclosed ranking rule.
  • Unsupported or irrelevant statistics reduce quality rather than improving it.

This is how “cutting-edge” advice remains correct after the headline has aged.

Publish limitations before someone else finds them

State:

  • What the evidence covers
  • What it does not cover
  • When it was checked
  • Which platform the recommendation applies to
  • Whether the conclusion is verified guidance, inference or experiment

Limitations do not weaken serious content. They define where it can be trusted.

Layer 5: become worth citing

Most content is technically acceptable and strategically replaceable.

If an answer engine already has 100 summaries of official documentation, summary number 101 has little reason to win. Google calls for non-commodity content: material with a perspective, experience or contribution that is not trivially reproduced.

Useful originality can take several forms:

  • First-hand implementation results
  • A transparent dataset
  • A decision framework
  • A comparison based on explicit criteria
  • A failure analysis
  • An expert interpretation of conflicting sources
  • A maintained reference that reconciles changes over time
  • A calculator, template, checklist or diagnostic tool

The Citation Readiness Model in this article is one example. It does not pretend to be a ranking factor; it converts fragmented platform guidance into an auditable workflow.

Add information gain, not word count

Before publishing, ask:

  • What does this page establish that the current top sources do not?
  • Which decision can a reader make after reading it?
  • Which claim comes from our own work?
  • Which claim has become outdated elsewhere?
  • What could a skeptical expert verify?

If the honest answer is “nothing,” adding another 1,000 words will not solve the problem.

What not to do in 2026

Several popular tactics are now either disproven, deprecated or routinely overstated.

Do not sell llms.txt as a Google ranking requirement

Google’s June 5, 2026 guidance explicitly says llms.txt and other special AI files are not needed for its generative Search features.

Maintaining an llms.txt file can still be an experiment or a convenient map for tools that voluntarily support the convention. Label it that way. Never let it replace crawlable HTML, internal links, sitemaps, canonical tags or standard access controls.

Do not promise FAQ rich results

Visible FAQs remain useful when they answer real objections or clarify scope. However, Google stopped showing FAQ rich results on May 7, 2026 and is removing related reporting and Rich Results Test support.

So publish FAQs for readers and topical completeness, not because someone promises extra Google result real estate.

Do not manufacture citations, reviews or mentions

Buying inauthentic mentions, adding irrelevant quotations or generating unsupported statistics may create the appearance of authority while making the page less reliable. Search spam systems and human reviewers are not the only risk: answer engines can associate the wrong claim with your brand.

Do not mass-produce prompt variations

One valuable guide with a clear decision path is stronger than 50 thin pages that swap synonyms. Scaled content without original value can violate spam policies regardless of whether AI helped create it.

Do not report visibility as revenue

A citation is not a visit. A visit is not a qualified lead. A lead is not a sale.

Keep the funnel honest.

A 90-minute citation-readiness audit

Use this sequence on one commercially or strategically important page.

Minutes 0–20: eligibility

  • Fetch the canonical URL and inspect its status, redirects and HTML.
  • Confirm indexability, snippet eligibility and canonical consistency.
  • Check Google Search Console URL Inspection and Bing URL Inspection.
  • Verify the sitemap contains only the preferred URL and an accurate lastmod.
  • Confirm important language alternatives are connected with reciprocal hreflang.

Minutes 20–40: interpretation

  • Compare the title, H1, description, byline and structured data.
  • Confirm the author links to a real profile with relevant credentials and identity references.
  • Check that the primary entity uses one name consistently.
  • Validate structured data and remove claims that are not visible on the page.

Minutes 40–60: extraction

  • Write down the five most important questions the page should resolve.
  • Find the exact passage answering each one.
  • Rewrite vague openings into complete answers with conditions.
  • Add descriptive headings where readers currently have to hunt.

Minutes 60–75: evidence

  • Mark every statistic, platform claim and recommendation.
  • Replace secondary summaries with primary evidence where available.
  • Add dates and scope to claims that can expire.
  • State uncertainty where the evidence is incomplete.

Minutes 75–90: differentiation

  • Identify what is genuinely original.
  • Add one useful artifact: a model, decision table, test, template or worked example.
  • Remove filler that merely restates the consensus.
  • Define the business action the article should support.

At the end, do not ask “Is this optimized for AI?” Ask: Is it eligible, intelligible, quotable, defensible and distinct?

A 30-day implementation roadmap

Days 1–3: establish the baseline

  • Export current organic landing-page data and conversions.
  • Record indexed URL counts and crawl errors.
  • Capture existing ChatGPT referrals using utm_source=chatgpt.com.
  • If available, record Bing AI Performance citations, cited pages and grounding queries.
  • Select five priority pages rather than rewriting the entire site.

Days 4–10: repair technical eligibility

  • Fix accidental crawler blocks, noindex directives and canonical conflicts.
  • Make primary content available in the initial rendered response.
  • Correct sitemaps and meaningful modification dates.
  • Verify language annotations.
  • Decide separately how to configure search crawlers and training crawlers.

Days 11–20: improve editorial evidence

  • Add direct answers to the priority questions.
  • Replace generic claims with bounded, sourced claims.
  • Strengthen author and publication information.
  • Add original examples, data or decision tools.
  • Remove stale sections and record substantive changes.

Days 21–25: publish and notify

  • Run schema, link, accessibility and performance validation.
  • Publish the updated canonical URLs.
  • Submit or refresh sitemaps.
  • Use IndexNow for the changed URLs where appropriate.
  • Request recrawling for a small number of critical Google URLs when justified.

Days 26–30: observe, do not overreact

  • Inspect crawler logs and indexing status.
  • Check whether grounding queries reveal missing subtopics.
  • Watch referrals and conversions.
  • Record citation observations with date, platform, prompt and location.
  • Avoid rewriting pages based on one unstable AI response.

Measure four different outcomes

A useful GEO/AEO dashboard separates four layers.

OutcomeEvidenceWhat it does not prove
DiscoveryCrawler logs, sitemap processing, URL inspectionThat the page was indexed or selected
Search visibilitySearch Console and Bing impressions, queries and positionsThat an AI answer used the page
Citation visibilityBing AI Performance and documented citation observationsThat users clicked or trusted the citation
Business impactReferrals, assisted conversions, leads and revenueThat the citation alone caused the outcome

Bing’s AI Performance public preview is particularly useful because it reports citation counts, cited pages and sampled grounding queries across supported Microsoft AI experiences. Bing warns that citation counts do not indicate ranking, authority or placement.

OpenAI provides a cleaner referral signal: links from ChatGPT search include utm_source=chatgpt.com. Track it, but remember that zero referrals does not prove zero citations; users do not always click.

For every manual citation test, record:

  • Platform and product
  • Account state or location where relevant
  • Exact prompt
  • Date and time
  • Whether web search was active
  • Cited URL
  • Whether the source materially supported the answer

Screenshots without this context are anecdotes, not measurement.

The next frontier: websites that agents can operate

Citation optimization concerns information retrieval. Agent readiness concerns task completion. They overlap, but they are not the same.

Modern browser agents may inspect screenshots, the DOM and the accessibility tree. Google’s web.dev guidance therefore recommends stable layouts, semantic structure and clear interactive states.

For sites where users ask agents to compare, book, buy or submit:

  • Use native HTML controls where possible.
  • Give form fields persistent labels.
  • Expose button names and states accurately.
  • Keep critical actions stable across layouts.
  • Avoid requiring hover-only discovery.
  • Make errors specific and programmatically associated with fields.
  • Test keyboard and screen-reader flows.
  • Require explicit confirmation for consequential actions.

This is not a magic GEO tactic. It is sound accessibility and interaction design that also makes delegated browsing more reliable.

A practical Citation Readiness Score

Score each layer from zero to five. This is an internal prioritization tool, not an engine score.

ScoreMeaning
0Missing or actively blocked
1Serious defects
2Basic implementation, inconsistent
3Reliable standard practice
4Strong, maintained implementation
5Distinctive, evidenced and continuously measured

Rate the page on:

  1. Discoverability
  2. Interpretability
  3. Extractability
  4. Verifiability
  5. Citation value

A page scoring 4 + 4 + 2 + 1 + 1 does not need another technical SEO sprint. It needs clearer answers, stronger evidence and a reason to exist.

That is the point of the model: direct investment toward the weakest gate.

The durable strategy

Nobody outside the platforms can promise a citation. The engines do not disclose complete selection systems, results vary by query and location, and the products continue to change.

What you can control is much more useful:

  • Whether the page is technically eligible
  • Whether its entities and purpose are clear
  • Whether important answers are easy to locate
  • Whether claims can be audited
  • Whether the work contributes something original
  • Whether changes are measured honestly

Do those things exceptionally well and you are not merely “optimizing for AI.” You are building a better source.

Frequently asked questions

Is GEO a replacement for SEO?

No. For Google, generative search remains grounded in the core Search index and ranking systems. GEO and AEO are useful labels for citation and answer visibility, but technical SEO, useful content and search eligibility remain the foundation.

Does structured data guarantee an AI citation?

No. Accurate structured data can clarify entities and make pages eligible for supported search features, but Google explicitly states that it neither guarantees rich results nor requires special schema for generative search.

Do I need an llms.txt file to appear in AI answers?

Not for Google Search. Google's June 2026 guidance explicitly says llms.txt and other special AI files are not required for its generative search features. An llms.txt file may still be maintained as an experimental convenience, but it should never replace crawlable HTML, sitemaps or standard metadata.

Should I allow GPTBot to improve visibility in ChatGPT Search?

GPTBot and OAI-SearchBot have different purposes. OAI-SearchBot controls eligibility for ChatGPT search answers, while GPTBot concerns potential training of OpenAI foundation models. A publisher can allow search crawling while disallowing training.

Does FAQ schema still produce Google FAQ rich results?

No for ordinary publishers. Google stopped showing FAQ rich results on May 7, 2026 and is removing related reporting and test support. Visible FAQs can still help readers and clarify a topic, but should not be presented as a current rich-result tactic.

How can I measure AI citation performance?

Combine Bing Webmaster Tools AI Performance, ChatGPT referral URLs containing utm_source=chatgpt.com, server logs, Search Console, analytics and conversion data. Treat citations, visits and conversions as separate metrics because an answer can cite a page without producing a click.

Sources

  1. Optimizing your website for generative AI features on Google Search · Google Search Central
  2. AI features and your website · Google Search Central
  3. Introduction to structured data markup in Google Search · Google Search Central
  4. Article structured data · Google Search Central
  5. Profile page structured data · Google Search Central
  6. FAQ structured data deprecation · Google Search Central
  7. Overview of OpenAI crawlers · OpenAI
  8. Publishers and developers FAQ · OpenAI
  9. Introducing AI Performance in Bing Webmaster Tools · Microsoft Bing
  10. IndexNow documentation · IndexNow
  11. GEO - Generative Engine Optimization · arXiv
  12. Build agent-friendly websites · web.dev

Further reading