AI SEO: how to get cited by ChatGPT, Claude, Perplexity & Google AI
Last updated 2026-06-10
AI assistants increasingly answer questions directly and cite a handful of sources instead of returning ten blue links. AI SEO — also called GEO (Generative Engine Optimization) or AEO (Answer Engine Optimization) — is how you become one of those cited sources. The mechanics are concrete and mostly under your control. This guide walks the four levers, in order of impact.
How AI answer engines pick which sources to cite
Whether it is ChatGPT search, Perplexity, Claude, or a Google AI Overview, the selection pipeline is roughly the same: a crawler indexes the open web, a retrieval step pulls candidate passages for a query, and the model synthesizes an answer and attributes the passages it leaned on. To be cited you must clear every stage — be crawlable, be retrievable (extractable + relevant), and be trustworthy enough to attribute. Miss any one and you are invisible no matter how good the content is.
Lever 1 — Let the right AI crawlers in
This is the prerequisite that everything else depends on. If the crawler can't fetch the page, none of the rest matters. Each major assistant uses distinct user-agents for distinct jobs — and the most common, costly mistake is blocking a search crawler when you meant to opt out of training. For example, you can allow OAI-SearchBot (ChatGPT search) and PerplexityBot (to be cited) while separately deciding on GPTBot (OpenAI training). Google splits this cleanly: Googlebot handles Search and AI Overviews, while the Google-Extended token controls Gemini training without touching your ranking.
See the full breakdown — user-agents, purpose, and copy-paste robots.txt rules — in the AI crawlers reference. If you want citations and AI-search traffic, allow at minimum the search/answer crawlers: OAI-SearchBot, PerplexityBot, Googlebot, and the on-demand fetchers ChatGPT-User + Claude-User.
Lever 2 — Be in the indexes AI engines actually retrieve from
Crawl access is necessary but not sufficient — each answer engine retrieves candidates from a specific index, and you have to be in it. The one most sites miss: ChatGPT search retrieves web results via Microsoft Bing's index(OpenAI's search partner), alongside its own OAI-SearchBot crawling. A page that ranks #1 on Google but is thin or absent in Bing is largely invisible to ChatGPT.
- ChatGPT → Bing's index. Check
site:yourdomain.comon bing.com, register at Bing Webmaster Tools (one-click import of your verified Google Search Console property), submit your sitemap. - Google AI Overviews / Gemini → Google's index. Standard Search Console + sitemap discipline covers you.
- Perplexity → its own index built by PerplexityBot — allowing the crawler (Lever 1) is the coverage step.
- Speed up both Bing and others with IndexNow pings on every publish — new URLs get indexed in hours instead of weeks.
Lever 3 — Make your content extractable
Retrieval favors content a machine can lift cleanly. Concretely: serve real HTML (most AI crawlers do not execute JavaScript, so client-only-rendered content is invisible to them); write quote-ready headings and a tight one-sentence answer near the top of each page; add schema.org structured data (Article, FAQPage, BreadcrumbList, and a Speakable block for voice); and keep one clear topic per URL. A short FAQ that mirrors real questions is one of the most reliably-cited formats because each answer is a self-contained, attributable passage.
Lever 4 — Be corroborated, fresh, and credible
Models prefer to cite claims they can corroborate. Facts that appear consistently across your site and independent sources are safer to attribute, so being referenced elsewhere compounds. Freshness matters too: keep a visible and structured dateModified, and update cornerstone pages on a cadence. And credibility — a clear author/organization identity, an about/methodology page, and outbound citations to primary sources — raises the odds a model trusts you enough to name you.
Where does llms.txt fit? (honest answer: optional)
llms.txt is a markdown file at your site root that lists the pages you most want cited, your crawler policy, and your attribution preferences. The honest 2026 status: no major AI provider has confirmed reading it, and Google has said it doesn't use it. It is not a citation lever today — the four levers above are. It IS a zero-cost forward bet on an emerging convention: if adoption comes, early publishers benefit; if not, you lost five minutes. If you publish one, keep it consistent with robots.txt — robots.txt is access control, llms.txt is stated intent, and they should never contradict each other. The free generator builds a spec-compliant file whose permitted/restricted sections mirror your robots.txt rules; the validatorscores any live site's file against the spec.
The practical AI SEO checklist
- Allow the search/answer crawlers you want citations from (see the crawler reference).
- Register in Bing Webmaster Tools + submit your sitemap — ChatGPT retrieves via Bing's index.
- Decide your training-crawler policy deliberately (GPTBot, ClaudeBot, CCBot) — separate from search.
- Serve real HTML; never gate primary content behind JavaScript.
- Lead each page with a tight, quote-ready answer; use clear H2s.
- Add Article + FAQPage + Speakable schema.
- Set and surface
dateModified; refresh cornerstone pages. - Establish author/organization identity + methodology for credibility.
- Optional: publish an llms.txt (cheap forward bet — see the honest status above), validate its score, and benchmark against the directory.
FAQ
What is AI SEO (GEO)?
AI SEO — also called GEO (Generative Engine Optimization) or AEO (Answer Engine Optimization) — is the practice of making your content easy for AI assistants and answer engines (ChatGPT, Claude, Perplexity, Google AI Overviews) to find, understand, and cite. Where classic SEO optimizes for a ranked list of blue links, AI SEO optimizes to be the source quoted inside an answer.
How do I get cited by ChatGPT?
Four things must be true: ChatGPT's crawlers can reach your page (allow OAI-SearchBot for search and ChatGPT-User for on-demand fetches in robots.txt), you are indexed in Bing (ChatGPT search retrieves results via Bing's index — register in Bing Webmaster Tools and submit your sitemap), your content is extractable (clean HTML, quote-ready headings, schema), and it is trustworthy and current. Being in the indexes is the prerequisite — blocking the crawlers removes you entirely.
Does AI SEO replace traditional SEO?
No — it extends it. The same fundamentals (crawlable, well-structured, authoritative, fresh content) feed both classic search and AI answer engines. AI SEO adds explicit AI-crawler policy in robots.txt, deliberate index coverage beyond Google (ChatGPT retrieves via Bing's index, so Bing Webmaster Tools suddenly matters again), and a stronger emphasis on machine-extractable structure.
Should I block AI crawlers to protect my content?
It is a real trade-off. Blocking training crawlers (GPTBot, ClaudeBot, CCBot) opts you out of model training but also reduces how well those models can describe and recommend you. Blocking search/answer crawlers (OAI-SearchBot, PerplexityBot) removes you from AI citations and the referral traffic they drive. Most publishers seeking visibility allow the search/answer crawlers and decide on training separately.
What is an llms.txt file and do I need one?
llms.txt is a markdown file at your site root that gives LLMs a curated map of your best content plus your AI-crawler policy. Honest status in 2026: no major AI provider has confirmed reading it, and Google has said it doesn't use it — so treat it as a zero-cost forward bet on an emerging convention, not a citation lever. The levers that demonstrably move citations are crawler access, Bing/Google index coverage, and extractable content. If you want one anyway, our free generator builds a spec-compliant file in under a minute.
How long does AI SEO take to work?
Like classic SEO, it compounds over weeks to months as crawlers re-index, your authority grows, and your content gets corroborated across sources. There is no instant switch — but allowing the right crawlers and shipping clean, extractable content is the prerequisite that everything else builds on.