AmarnepalNepal Data
AEO & GEOAdvanced · 9 min read · verified 2026-07-28

llms.txt and AI crawlers

AI crawlers are controlled separately from Googlebot — GPTBot and OAI-SearchBot for OpenAI, Google-Extended for Gemini. Deciding what to allow is a business decision, not a default. llms.txt is a proposed convention for pointing LLMs at your best content, and it is a cheap experiment rather than a strategy.

If you want AI answer engines to cite you, two things have to be true: they must be permitted to fetch your content, and what they fetch must be worth quoting. This guide covers the first properly and the second briefly.

The permissions question is genuinely under-examined. Many sites have robots.txt rules written years ago for search crawlers, which now silently determine whether they appear in AI answers at all — in either direction.

The crawlers that matter, and what each does

AI companies run separate crawlers for separate purposes, and conflating them leads to bad decisions. The important distinction is between crawling to train models and crawling to answer a live user query with attribution.

  • GPTBot (OpenAI) — crawls to improve models. Documented publicly with its user agent and IP ranges.
  • OAI-SearchBot (OpenAI) — crawls for search and live answers, which is the one associated with being surfaced and linked in ChatGPT search results.
  • Google-Extended — not a crawler in itself but a robots.txt token controlling whether content Googlebot fetches may be used for Gemini and related generative products. Independent of Search ranking.
  • Googlebot — classic search crawling. Blocking it removes you from Google Search; it is not an AI control.
  • Other engines run their own agents, and the list changes. Check current vendor documentation rather than trusting a blog post's list — including this one.

Deciding what to allow

This is a business decision with a genuine trade-off, and reasonable publishers land in different places.

Allowing search-and-answer crawlers is how you become citable in AI answers, with attribution and potential referral traffic. Allowing training crawlers contributes your content to model capabilities with no attribution and no referral. Many publishers allow the former and block the latter, which is a coherent position.

Blocking everything is also coherent if your content is your product. What is not coherent is having no position and discovering by accident that a decade-old robots.txt is making the choice for you.

  • Decide separately for training crawlers and for search/answer crawlers — they are different bargains.
  • Note that Google-Extended is independent of Search: blocking it does not affect your Google ranking.
  • Check your current robots.txt actually expresses your intention, then re-check after any platform migration.
  • Verify with server logs that the agents you expect are actually fetching you.

What llms.txt is

llms.txt is a proposed convention: a markdown file at the root of your domain that gives LLM consumers a curated, readable map of your most useful content, avoiding the navigation, scripts and boilerplate of rendered HTML.

The idea is reasonable — the same instinct as robots.txt and sitemap.xml, applied to a new consumer. The honest status is that it is a proposal rather than an adopted standard, and support among major AI vendors remains limited.

  • A short description of what the site is and who publishes it.
  • Curated links to your genuinely most useful pages, grouped by topic, each with a one-line description.
  • Pointers to canonical data or reference pages you would want quoted.
  • Kept short and curated. A dump of every URL defeats the purpose — that is what a sitemap is for.

An honest assessment of llms.txt

It costs an hour to write and essentially nothing to host, and it may help. That is the entire case for it, and it is enough.

What it is not: a ranking mechanism, a guarantee of citation, or a substitute for being crawlable, accurate and well-structured. If someone is selling llms.txt authoring as a GEO service, that is the tell.

Publish one if you like, and then spend the remaining effort on the things that demonstrably matter — specificity, sourcing, dating, structure and authority.

Writing content AI can quote confidently

Permissions get you fetched; this gets you cited. The constraint on a generative system is that it must not assert something false and attribute it to you, so it favours what it can state safely.

  • Give concrete figures with units, not hedged ranges — and say when you verified them.
  • Name your sources explicitly and link them, preferring primary ones.
  • Show the author and the review date on the page, and in Article markup.
  • State uncertainty where it exists. 'Reported at both Rs 2,000 and Rs 3,000; confirm at the counter' is more quotable and more honest than inventing precision.
  • Structure for chunk retrieval: question-shaped headings, self-contained answers, real lists and tables, a summary at the top and a takeaways block.
  • Keep your entity identity consistent everywhere, and use Organization markup with sameAs links.

Checking whether any of it worked

Measurement here is genuinely poor and it is worth saying so rather than pretending otherwise.

The practical routine: check your server logs for AI crawler user agents to confirm you are being fetched at all; query your key topics in ChatGPT, Perplexity, Gemini and Google AI Overviews on a schedule and record whether you are cited; and watch analytics for referrals from AI hostnames, accepting that attribution is incomplete.

Set a monthly cadence, keep a simple record, and treat any tool promising precise AI-citation analytics or guaranteed placement with the scepticism it deserves.

Key takeaways

  • AI crawlers are controlled separately from Googlebot — GPTBot, OAI-SearchBot and the Google-Extended token.
  • Decide separately about training crawlers and search/answer crawlers; they are different bargains.
  • Blocking Google-Extended does not affect your Google Search ranking.
  • llms.txt is a proposal with limited vendor adoption — a cheap experiment, never a strategy.
  • Citation comes from dated, sourced, specific facts in a chunk-retrievable structure, not from a config file.
Questions

llms.txt and AI Crawlers — FAQ

What is llms.txt?+

A proposed convention: a markdown file at your domain root giving LLM consumers a curated, readable map of your most useful content, without the navigation and boilerplate of rendered HTML. It is a proposal rather than an adopted standard, and support among major AI vendors is limited.

Should I block AI crawlers?+

It is a business decision with a real trade-off, and it should be made separately for training crawlers and for search/answer crawlers. Search and answer crawlers are how you become citable with attribution and referral traffic; training crawlers offer neither. Many publishers allow the former and block the latter.

What is the difference between GPTBot and OAI-SearchBot?+

OpenAI documents GPTBot as crawling to improve its models, and OAI-SearchBot as crawling for search and live answers — the one associated with being surfaced and linked in ChatGPT search results. They are controlled separately in robots.txt.

Does blocking Google-Extended hurt my Google rankings?+

No. Google-Extended is a separate robots.txt token controlling whether your content may be used for Gemini and related generative products, and it is independent of Google Search crawling and ranking.

Will llms.txt get me cited by ChatGPT?+

There is no evidence it reliably does, and vendor adoption is limited. It is cheap enough to be worth publishing as an experiment, but citation comes from being crawlable, specific, dated, sourced and well-structured — not from a configuration file.

How do I know if AI crawlers are visiting my site?+

Check your server access logs for the documented user-agent strings, which the vendors publish. That confirms fetching. Whether you are actually cited is separate and must be checked by querying the engines directly, since no vendor reports citation frequency.

Related guides

← All guides

Sources & data note

This guide describes documented, widely-accepted practice as published by Google Search Central, web.dev and schema.org, which are cited above. Search and AI systems change continually: treat specific thresholds and crawler names as current guidance and verify against the official documentation before relying on them. Nepal-specific observations — market conditions, Devanagari search behaviour, what local competitors do — are our own analysis rather than published findings, and are not separately sourced. Guides are written from primary sources — Nepali government departments, operators, park authorities and standards bodies — and each guide lists the sources used for its own facts. Rules, fees and prices in Nepal change; treat figures as current at the review date shown on each guide and verify anything money- or visa-critical with the issuing authority before you rely on it.