Guide9 min readPublished 2026-10-10

llms.txt best practices

Practical llms.txt best practices: what to publish, how to structure it, and what it can and cannot do for AI visibility.

OF
OmniForce
Product Team
#geo#llms-txt

Direct answer

Publish one markdown file at /llms.txt: an H1 with your site name, a one-sentence blockquote summary, then H2 sections of curated links with short descriptions. Keep the main file to a few thousand words, use absolute URLs, and treat it as a routing map for agents passing through — never as a ranking lever.

What llms.txt is actually for

Start with what it isn't: a submission mechanism. There is no registry, no queue, no ping endpoint. The convention — an H1, a blockquote, then link lists grouped under H2s — was proposed in September 2024 and lives in a GitHub repo that has stayed deliberately small.

The mechanism is retrieval, not indexing. By the time anything reads your file, a decision has already been made to fetch your site: a browsing assistant, a coding agent, a RAG pipeline with your domain on its allowlist. That reader arrives with a task and a token budget. /llms.txt is the map you hand it at the door.

Google's guidance on AI features in Search is blunt about the indexing half of the bargain: no special file or markup is required for your pages to appear in AI-powered results. Crawler access is a separate question, governed by robots.txt, and OpenAI documents exactly how that works for its crawlers. You cannot opt into AI visibility with a text file any more than you can opt into a search index with a sitemap alone.

Picture a press kit on the reception desk. It tells a visitor where the archive is and who to ask for. It does not decide who gets through the door.

The anatomy of a file that gets read

Most of the files we inspect are link dumps — sixty URLs, no descriptions, no order. A reader skims them for a domain match and moves on. Files that actually get used share six properties.

One H1, matching the name people search for. A blockquote of a single sentence that says what the thing is and who it serves. Two or three lines of plain context — base URL, auth model, what lives where. H2 sections named for jobs rather than departments, so a reader looking for "how do I cancel a subscription" finds a heading it recognizes. Descriptions under ten words, sitting next to each link. And an Optional section at the bottom for the secondary material that would otherwise bloat the top.

Here is what that looks like in practice.

# Northwind Analytics

> Northwind is a billing API for subscription software teams. Docs, SDKs, and pricing live here.

Northwind exposes a REST API, four official SDKs, and a hosted webhook relay.
Base URL: https://api.northwind.example/v2

## Core documentation
- [Quickstart](https://docs.northwind.example/quickstart): auth, first charge, five minutes
- [API reference](https://docs.northwind.example/api): every endpoint, request and response
- [Webhooks](https://docs.northwind.example/webhooks): events, retries, signature checks

## SDKs
- [Python](https://docs.northwind.example/sdk/python)
- [TypeScript](https://docs.northwind.example/sdk/typescript)

## Pricing and limits
- [Plans](https://northwind.example/pricing): tiers, overage, annual discounts
- [Rate limits](https://docs.northwind.example/limits): per-key quotas and backoff

## Optional
- [Changelog](https://northwind.example/changelog)
- [Status](https://status.northwind.example)

Two details do most of the work. Absolute URLs, because a reader may be parsing your file out of context and has no reliable base to resolve against. And short descriptions, because they are the only signal that tells a reader which of six links under "Core documentation" answers the question it is holding.

There is a useful precedent for this genre in Hugging Face model cards: a short block of structured prose that humans skim and machines parse, sitting next to something far larger. Your llms.txt is a model card for a website.

Hand-maintained files rot. Endpoints deprecate, pricing pages move, a product gets renamed and the map keeps pointing at the old address. Pull the file from your sitemap and your frontmatter at build time — llms.txt generation wired into the same pipeline that builds your docs means the map cannot drift more than one deploy behind the territory.

Agent traffic is where it pays off

Chatbots answering a question are the least demanding readers you will get. They fetch one page and synthesize. Agents are different: they hold a goal across many steps, choose tools mid-task, and MCP gives them a standard way to call those tools. An agent that has just been asked to "upgrade this customer's plan" needs to find your endpoint, your auth scheme, and your error codes — in that order, without fetching twelve pages to learn which one matters.

This is the tradeoff that decides your file's shape. Cover everything, and you spend context you don't get back. llms-full.txt exists for the case where a reader genuinely wants the corpus in one shot, and for a documentation site that can be a 40,000-word payload that gets truncated before the useful section arrives. Our default: a tight llms.txt as the map, and llms-full.txt for the readers that ask for it by name.

Concretely, a support agent resolving a billing question reads your file, sees [API reference](...), fetches that one page, finds the endpoint, and stops. Three requests instead of a crawl. That is the whole return on the effort.

What it cannot do

It cannot control crawlers. That is robots.txt, and Google's introduction to the protocol is still the right place to learn what a user-agent directive will and won't accomplish. A file that says "please cite us" while your robots.txt blocks the crawler is a note left in a locked room.

It cannot establish who you are. Entity identity comes from structured data — the Organization, Product, and article types at schema.org — which give a machine a canonical name, logo, and sameAs relationships. Agent-ready structured data is what lets a model resolve your brand to a real entity instead of a string that appears on eleven unrelated pages.

It cannot fix thin content. If your docs sit behind a login, or your best answer lives in a PDF from 2019, a tidy map of links leads nowhere. Bing's search team publishes webmaster guidance on how Copilot surfaces content, and the answer there rhymes with Google's: crawlable, well-structured pages, described plainly. No file shortcuts that.

What it can do is small, real, and cheap: shave requests off a retrieval loop for readers who were already coming.

Measuring, then maintaining

Your server logs are the only hard evidence. Count GET /llms.txt, look at the user agents, and watch the trend over a quarter. Everything past that is inference. There is no public data on what share of sites publish the file or which engines consume it, so anyone quoting a lift figure is quoting something they made up.

The maintenance loop we run is five steps, and only the first is about the file itself.

  1. Log the fetches. If nothing has requested the file in ninety days, the file is not your bottleneck.
  2. Track citations. Answer-engine citation tracking tells you which of your URLs are being quoted, which is a different list from the ones you think are important.
  3. Find the decay. Pages that lose citations over time are usually pages that lost accuracy — a price change, a deprecated endpoint, a version bump. Content decay detection and refresh catches that before a competitor's page replaces yours as the source.
  4. Fix the gaps. Run a GEO / AI-visibility audit across your domain to see which questions you are absent from, then check whether the page that should answer them exists, is crawlable, and is described in your file.
  5. Regenerate on deploy. The file is an artifact, not a document.

The order matters. Teams that start at step five spend a quarter writing a beautiful map to a site that answers nothing.

FAQ

Does llms.txt improve my rankings in AI Overviews or ChatGPT?

No. Google's AI-features guidance states that no special file or markup is needed for pages to appear, and the file plays no role in ranking. Treat it purely as a retrieval aid for readers who have already arrived.

Do OpenAI, Anthropic, or Google actually read it?

Neither OpenAI's crawler documentation nor Google's AI-features guidance names the file as an input. Anthropic publishes one of its own for its developer docs, which suggests the format is understandable to the people who build these systems. Understandable is not the same as consumed, and the honest position is that support varies and is not published.

Should I ship llms.txt, llms-full.txt, or both?

Both is a defensible default. The short file serves agents that need to navigate; the full file serves pipelines that want the corpus in a single fetch. If you can only maintain one, maintain the short one — a stale map is survivable, stale content is not.

Where does the file live if my docs are on a subdomain?

At the root of the host that serves the content, so docs.example.com/llms.txt covers the docs. A root-domain file works as an index to subdomain files if you have several. Absolute URLs in the body remove most of the ambiguity either way.

How often should I update it?

Whenever the underlying navigation changes — a new section, a retired product, a moved pricing page. In practice that means every deploy, generated automatically, with a human reviewing the diff when the structure changes rather than when a paragraph does.

Next step

Before you write a line of the file, find out what the engines already say about you. A GEO audit shows which of your pages get cited, which questions you're missing from, and whether a map would help anyone navigate — or just describe an empty building.

Run a free GEO audit of your site

Related reading

OmniForce Engine

See how AI engines read your site

Run a free audit to check your structured data, machine-readable files, and crawler access.