How to get cited by AI assistants (ChatGPT, Perplexity, Claude, Gemini)
Getting cited by an AI assistant comes down to four things, in order. First, be reachable: if your robots.txt blocks the search and citation crawlers, you cannot be cited, full stop. Second, be liftable: lead each page with a direct, self-contained answer the model can quote without untangling it from preamble. Third, be parseable: clear headings, schema markup and an honest FAQ let a model read and quote you with confidence. Fourth, be trusted: earn mentions and citations on independent sites the engines already read, because a reference from a source a model trusts carries more weight than another link pointing only at your own domain. None of this is a trick or a hack. It is the same discipline as good content work, aimed at a new reader: a model composing one answer on behalf of a person about to act on it. This guide walks each step, shows what to check, and explains how to know whether it worked across the engines that matter.
Being cited by an AI assistant is not luck, and it is not a growth hack. It is the predictable result of four things done in order: be reachable, be liftable, be parseable, be trusted. Skip the first and the rest cannot help you. Do all four and citations stop being a surprise.
This is the practical version. No borrowed statistics, no secret tactic. Where measurement matters, we point to our own ongoing tracking rather than cite a number we cannot verify.
Step 1: be reachable
If the engines cannot fetch your page, nothing else matters. The most common self-inflicted wound in GEO is a robots.txt that blocks the very crawlers that produce citations.
Here is the part most people miss: every provider runs two separate crawlers, a training crawler and a search or citation crawler, and they are controlled by different tokens. Blocking the training crawler is a defensible choice. Blocking the search or citation crawler removes you from the live answers you want to win.
| Provider | Search / citation crawler (allow for GEO) | Training crawler (decide separately) |
|---|---|---|
| OpenAI | OAI-SearchBot, ChatGPT-User | GPTBot |
| Anthropic | Claude-SearchBot, Claude-User | ClaudeBot |
| Perplexity | PerplexityBot, Perplexity-User | (none distinct) |
| (Gemini grounding via search) | Google-Extended |
Allow the search and citation crawlers. Two notes from how these actually behave: the tokens are matched case-insensitively as substrings, and at least one engine takes around a day to propagate a robots change, so do not judge the effect the same afternoon.
Step 2: be liftable
A model composes an answer by lifting clean, self-contained statements from sources it trusts. If your answer is tangled inside three paragraphs of introduction, it gets skipped in favor of a page that states it plainly.
So write answer-first. Lead the page, and ideally each section, with the direct answer in one or two sentences, then expand underneath. A reader who wants depth keeps going; a model that wants a quotable answer already has one. This single habit tends to move citations more than any other on-page change.
Keep claims specific and checkable. "Faster onboarding" is weak. "Onboarding in one form, with the first audit back the same day" is liftable, because it is concrete and a model can quote it without hedging.
Step 3: be parseable
Reachable and liftable get you most of the way. Structure gets you the rest, by letting a model read your page with confidence about what each part means.
- Use a clear heading for each question or topic, so the structure of the page mirrors the structure of the answers people ask for.
- Add schema markup: Article, FAQPage and Organization at minimum. It tells the model what it is reading rather than making it guess.
- Write a real FAQ that answers the questions people actually ask, in their words, with direct answers. An honest FAQ is some of the most liftable content on a site.
- Keep one clear topic per page. A page that tries to answer everything answers nothing cleanly.
Structure is not decoration. It is the difference between a model quoting you confidently and a model deciding your page is too ambiguous to risk.
Step 4: be trusted
The last step is the hardest and the most durable. Models lean on sources they already trust, and they read your brand across the whole web, not just on your own domain. A mention or citation on an independent site the engines respect carries more weight than another backlink pointing only at you.
That means earning genuine references: being included in a roundup a model reads, answering in communities the engines index, being the source other people cite when they explain your category. This is slow, it cannot be faked convincingly, and it is exactly why it works. Being referenced reads, to a model, as being trusted.
How to know it worked
You cannot manage what you do not measure, and ranking position is the wrong metric here. Track presence in answers instead:
- Mention rate: how often the engines name you for the prompts that matter.
- Citation rate: how often they link you as a source, not just mention you.
- AI share of voice: your slice of the answers versus your competitors.
- Position in the answer: where you sit in the reading order, not a search rank.
Check these across engines, because presence on ChatGPT does not imply presence on Perplexity, Claude or Gemini. Then treat the whole thing as a loop: publish, verify, adjust, repeat. That continuous loop, run across every engine, week after week, is where DIY efforts usually run out of road, and it is exactly the part Surface Agent runs for you.
Put the four steps to work for you
Doing all four, week after week and across every engine, is the job Surface Agent handles for you. It watches your visibility across ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews around the clock, along with the Reddit threads where your buyers ask for recommendations. It drafts the answer-first pages and structure the engines reward, drafts the Reddit replies in your voice for you to post from your own account, and drafts the media outreach that earns the third-party references. You review and approve, and the citations compound.
Frequently asked questions
01Why am I not being cited by ChatGPT or Perplexity?
The most common reasons, in order: your robots.txt blocks the AI search and citation crawlers; your pages bury the answer under preamble so the model cannot lift it cleanly; your content lacks the structure (headings, schema, FAQ) that lets a model quote you confidently; or you have little third-party presence on sources the engines already trust. Work through them in that order.
02Does blocking GPTBot or ClaudeBot stop me from being cited?
Be precise about which bot. The training crawlers and the search or citation crawlers are separate tokens controlled independently. Blocking a search or citation crawler (for example OAI-SearchBot or Claude-SearchBot) removes you from those answers. If visibility in AI answers is the goal, allow the search and citation crawlers; decide on training crawlers separately.
03How long until citations show up?
It varies by engine and by how often each recrawls. Crawler directive changes can take roughly a day to propagate on some engines, and earned references take longer to be discovered and reflected. Treat it as a continuous loop, not a one-time fix: publish, verify whether the engines now cite you, adjust, repeat.
04What is the single highest-impact change?
For most sites that are otherwise reachable, it is answering the real question directly at the top of the page in plain, self-contained language. Models lift clean answers; they skip pages where the answer is tangled in introduction. Reachability is the prerequisite, but answer-first writing is usually the change that moves citations the most.