Search "AI search optimization" and you'll find a stack of agency blog posts telling you to restructure your content into "AI-friendly chunks," publish an llms.txt file, and add extra schema markup so ChatGPT "trusts" your page. Almost none of it links to a primary source, and some of it directly contradicts what the AI operators themselves have published. This article does the opposite: every claim below traces to a named operator's own documentation, dated, with a link. Where the evidence stops, it says so instead of filling the gap with something that sounds plausible.
The AI crawlers, named correctly
The first mistake in most AI search optimization advice is treating "AI bots" as one category. OpenAI alone runs four crawlers with four jobs: GPTBot trains models, OAI-SearchBot indexes for ChatGPT's search results, ChatGPT-User fetches a page only when a prompt triggers it (not a background crawl), and OAI-AdsBot checks landing pages for ad policy, per OpenAI's own bot documentation. You can block training without blocking search visibility, just by writing separate robots.txt rules for each agent.
Anthropic reorganized its docs the same way. As of an April 7, 2026 update, it names three bots: ClaudeBot for training, Claude-User for a live fetch tied to one question, and Claude-SearchBot for indexing content that feeds search-style answers, distinct from training. Anthropic's support page gives the names and robots.txt tokens but doesn't publish a literal user-agent string. A post printing an exact "ClaudeBot/1.0" string as confirmed is going beyond what Anthropic itself has published; treat the names and tokens as solid, the UA string as unverified. Search Engine Land and Search Engine Journal both covered the split.
Perplexity runs two: PerplexityBot indexes for search, and Perplexity-User fetches live on a user's behalf. Per Perplexity's own docs, Perplexity-User "generally ignores robots.txt rules" for that fetch. Google-Extended and Applebot-Extended work differently: neither has a separate user-agent, both are robots.txt-only tokens. Google's crawler documentation, updated July 14, 2026, says blocking Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" there. It only opts out of Gemini/Vertex training and grounding. Apple's page on Applebot-Extended, dated June 8, 2026, says the same for Siri, Spotlight and Safari.
| Bot | Operator | Purpose | robots.txt token |
|---|---|---|---|
| GPTBot | OpenAI | Training | GPTBot |
| OAI-SearchBot | OpenAI | ChatGPT search indexing | OAI-SearchBot |
| ChatGPT-User | OpenAI | Live fetch on user request | ChatGPT-User |
| OAI-AdsBot | OpenAI | Ad-page policy checks | OAI-AdsBot |
| ClaudeBot | Anthropic | Training | ClaudeBot |
| Claude-SearchBot | Anthropic | Search-style answer indexing | Claude-SearchBot |
| Claude-User | Anthropic | Live fetch on user request | Claude-User |
| PerplexityBot | Perplexity | Search indexing | PerplexityBot |
| Perplexity-User | Perplexity | Live fetch, often ignores robots.txt | Perplexity-User |
| Google-Extended | Gemini/Vertex training + grounding opt-out | Google-Extended | |
| Applebot-Extended | Apple | Apple Intelligence training opt-out | Applebot-Extended |
Wanting ChatGPT search visibility but not training inclusion takes two rules, not one: Disallow for GPTBot, nothing for OAI-SearchBot. Get that backwards and you block the bot you wanted. Check your robots.txt against the table above with the Robots.txt Tester first.
Does llms.txt actually work
Short answer, sourced: no confirmed evidence that it does anything for AI search visibility, and one direct statement from Google that it isn't needed.
No operator has confirmed reading llms.txt for search
Google's John Mueller answered a question about it directly on Reddit's r/TechSEO: "Google doesn't use llms.txt or llms-author.txt. I don't know of any other crawler / llm confirming they're using these (other than SEO tools)." Search Engine Journal reported the exchange on July 6, 2026.
That lines up with Google's own guidance: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search," with llms.txt named explicitly among the things you can skip, per Google's AI optimization guide, updated July 10, 2026.
Ahrefs put a number on actual usage: of 137,210 domains studied, roughly 28% publish an llms.txt file, and 97% of those got zero requests for it in May 2026. The small remainder that did get hit mostly saw SEO audit tools and general crawlers, not the operator-named bots above. Ahrefs' own sample skews toward technical, SEO-aware sites, meaning real web-wide adoption is probably lower still. That's not a standard being adopted. That's a file most sites publish once and nobody, human or bot, ever comes back to read. Full study on Ahrefs' blog, published June 15, 2026.
One nuance worth flagging: Anthropic's engineering team has written about pointing coding agents at a software library's llms.txt file so Claude Code reads docs faster. That's real, and unrelated to a chatbot deciding whether to cite your blog post. Conflating the two is how "Anthropic recommends llms.txt" turns into a claim with no source. Nobody, Google, Anthropic, or Perplexity, has published guidance saying llms.txt affects search citation. Wanting one for the coding-agent case is defensible. "It'll get me cited" currently isn't.
Structured data and AI Overviews: what Google actually said
Structured data is the other place where confident claims outrun the sourcing. Google has said, twice, in two separate documents in the last year, that it isn't required for AI-generated results. The AI optimization guide states: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add." The AI features documentation, updated December 10, 2025, repeats it almost verbatim: "There's also no special schema.org structured data that you need to add."
That doesn't make schema markup pointless. It remains the established mechanism for classic rich-result eligibility, a separate benefit from AI citation, and worth keeping for that reason alone. A product page carrying flawless Product schema still won't show up in an AI Overview if the site never opted into "generative AI features" in Search Console; the markup and the eligibility toggle solve two different problems.
| Claim | Status |
|---|---|
| Structured data required for AI-generated results | Contradicted directly by Google |
| Schema markup improves classic rich-result eligibility | Confirmed, separate mechanism from AI features |
| Schema markup reduces "AI confusion" or acts as a citation trust signal | Unsourced, no Google statement or named study found |
| A specific percentage lift in citation odds from adding schema | Unsourced, circulates in SEO blog posts without attribution |
| Indexing + snippet eligibility + AI-features opt-in required for AI Overviews | Confirmed by Google, "not guaranteed" even then |
The middle rows are where AI search optimization content quietly stops being sourced: claims that schema "reduces AI confusion" or acts as a "trust signal" circulate with no attribution to any Google statement or study, which makes them unproven, not necessarily false. A precise percentage lift tends to show up anyway, because a made-up number reads more convincingly than an honest shrug.
What to actually do
Skip the speculative fixes and work the confirmed list instead.
Get the basics right first: your page must be indexed and snippet-eligible in classic Search before it can show up in any AI-generated result, since AI Overviews draw from the same index rather than a separate quality bar. Check the "generative AI features" toggle in Search Console; if it's off, you're opted out entirely. The AI Readiness Checker audits a page against these confirmed factors instead of someone else's checklist.
Keep the content that matters in actual text, not locked inside an image or video with no transcript, per Google's AI features documentation. Don't rewrite pages into short, fragmented chunks on the theory that AI systems need bite-sized pieces. Google's guidance says the opposite: its systems parse a full page and pull the relevant section themselves. Structure the page for the humans reading it.
Check what each AI crawler can reach with the Robots.txt Tester, using the operator-confirmed token names above rather than one generic "AI bot" rule.
If you still want an llms.txt file, be honest about why. "It'll help with search citation" doesn't hold up against the Ahrefs numbers and Mueller's statement above. "Low-cost, might help a coding agent" is a real, narrower reason. The llms.txt Generator builds the file in a couple of minutes either way.
