llms.txt: Where It Goes, Who Reads It, and How to Publish One
What is llms.txt?
The important distinction is function: llms.txt curates context; it does not replace robots.txt for crawler access or sitemap.xml for URL discovery. The proposal was introduced by Jeremy Howard of Answer.AI in September 2024 and updated to v2 in August 2026. A file can live at /llms.txt for the whole site or at a subpath such as /docs/llms.txt for pages beneath that path. It can include background, guidance, and curated links to detailed pages or Markdown twins.
That makes llms.txt a curation layer, not a control file. It does not grant crawler permission, force indexing, verify a claim, or guarantee a citation.
| File | Documented role | What it does not prove | Format |
|---|---|---|---|
robots.txt | Communicates crawler access rules by user agent | That an allowed crawler will index or cite a page | Plain-text directives |
sitemap.xml | Lists canonical URLs for search-engine discovery | That listed URLs will rank | XML |
llms.txt | Proposes curated context and links for agents | That a platform consumes, weights, or cites it | Markdown |
Key Takeaway: llms.txt complements crawler access and search discovery; it does not replace either one.
For the broader discipline, see the LLM SEO guide, the GEO guide, and the AEO vs GEO comparison.
Is llms.txt a feed?
No. It is not a syndication or feed protocol like RSS or Atom, and it does not push updates to anyone.
Is llms.txt a consent or permissions file?
No. It does not grant or refuse permission to crawl or use content. Platform-specific crawler controls remain the documented way to manage access.
Does ChatGPT Use llms.txt?
OpenAI publishes llms.txt files for parts of its own developer documentation, but its crawler documentation does not say that ChatGPT Search consumes third-party llms.txt files as a ranking or citation signal. Those are different facts: self-publication demonstrates producer adoption of the format; it does not document how OpenAI selects external sources.
OpenAI documents OAI-SearchBot, GPTBot, and ChatGPT-User and explains how site owners can control them through robots.txt. A request to your llms.txt from one of those user agents is evidence of access only. It does not prove the file was parsed, weighted, or responsible for a later citation.
Evidence status: Unproven. There is no OpenAI documentation establishing that ChatGPT reads third-party site llms.txt files as a ranking or citation signal, and crawler access documentation is not proof of llms.txt consumption.
Key Takeaway: For ChatGPT, third-party llms.txt consumption remains unproven in public documentation. Publish the file for optional machine-readable context, not as a guaranteed ChatGPT lever.
Does Google Use llms.txt? (The Honest Answer)
Google Search does not require a special AI-only file for AI Overviews or AI Mode. Google's official guidance for generative AI features in Search says the established SEO foundations still apply: pages must be crawlable, indexable, useful, and eligible to appear in Search. That guidance does not document llms.txt as a ranking or citation signal for Google Search.
This conclusion is intentionally scoped to Google Search and its generative Search features. It is not a claim about every Google product, browser feature, developer tool, or future experiment. A Google crawler requesting a file also does not establish that Search uses its contents.
Key Takeaway: Use ordinary technical SEO and useful, verifiable content for Google Search, AI Overviews, and AI Mode. Do not sell llms.txt as a Google ranking shortcut.
The underlying authority work is explained in the Reference Authority framework.
Why Might an llms.txt File Still Be Worth Publishing?
llms.txt matters as a low-cost experiment in machine-readable curation, not as a proven visibility factor. Web pages contain navigation, scripts, repeated templates, and other material an agent may not need. A short manifest can offer a clean site description and a curated route to authoritative pages when an agent or custom retrieval workflow is designed to use it.
It is worth prioritizing when a site has substantial documentation, a complex information architecture, stable canonical resources, or an audience that uses coding agents and custom RAG tools. It is lower priority when the site still has blocked crawlers, weak indexability, inconsistent entities, thin public evidence, broken canonical URLs, or no maintained source pages. Fix those foundations first.
The practical upside is controlled curation and easier machine access. The limitation is equally important: publication alone does not show that a commercial AI platform discovered the file, trusted the descriptions, or changed an answer because of it.
Key Takeaway: Publish llms.txt when the curation is useful and maintenance is affordable. Do not let it displace crawlability, content quality, entity clarity, or earned authority.
The Exact llms.txt Format (Per llmstxt.org)
The August 2026 v2 proposal keeps an H1 as the only required element and allows supporting context plus H2-delimited file lists. A conventional production file contains:
- H1 name. The project or site name; this is the only required element.
- Blockquote summary. A concise description that helps an agent interpret the file.
- Context or guidance. Optional Markdown below the summary, without another heading before the file-list sections.
- H2 sections. Categories such as
## Docs,## Guides, or## Services. - Markdown link lists. Each item requires a Markdown link. Notes after the link are optional, although specific descriptions make a curated file more useful.
- Optional secondary grouping. A section named
## Optionalmay still organize lower-priority links, but v2 no longer gives that name a mechanical omit-from-context meaning.
A root file can cover the whole site. V2 also defines path-scoped files: /docs/llms.txt covers pages beneath /docs/, and the most specific applicable file takes precedence. Pages may expose clean Markdown twins by appending or replacing the extension with .md. V2 also proposes discovery through rel="alternate" type="text/markdown" for a page's Markdown version and rel="describedby" for the applicable llms.txt, supplied through HTML links or HTTP Link headers.
A compact example:
# Acme Invoicing
> Automated invoice reconciliation for finance teams.
Acme publishes public product, security, and API documentation.
## Docs
- [Getting Started](https://acme.com/docs/getting-started): Setup guide for new accounts.
- [API Reference](https://acme.com/docs/api): REST API and webhook reference.
## Pages
- [Security](https://acme.com/security): Public security and data-handling policies.
## Optional
- [Changelog](https://acme.com/changelog): Product updates.The ## Optional section in this sample remains permitted by the proposal, and its own example still uses it, but it is not required and has no documented ranking meaning. Use fully qualified canonical URLs and accurate notes. The proposal permits more Markdown than older summaries implied; the goal is predictable, useful context rather than invented parser rules.
Key Takeaway: V2 is flexible but clear: one required H1, optional context, H2 file-list sections, optional notes, path scope, and discoverable Markdown twins.
What Do We Know About How AI Tools Use llms.txt?
Current public evidence supports four different statuses, and they should not be collapsed into one claim.
| Status | What it means | Current example |
|---|---|---|
| Documented | A primary specification or vendor document states the behavior. | llmstxt.org defines v2; vendors document crawler controls through robots.txt. |
| Observed | A file or request can be seen, but its downstream use is not established. | OpenAI, Anthropic, Perplexity, and other documentation sites publish llms.txt files. |
| Plausible | A custom agent or RAG workflow can be configured to discover and use the format. | Documentation assistants can follow a curated manifest and Markdown twins. |
| Unproven | No primary documentation establishes the claimed ranking, weighting, or citation effect. | Third-party llms.txt consumption by major consumer AI search products. |
OpenAI: publishes the format for developer documentation and documents web crawlers separately. Anthropic: documents ClaudeBot, Claude-User, and Claude-SearchBot through robots.txt, but that crawler guidance does not document third-party llms.txt consumption. Perplexity: publishes an llms.txt for its docs and documents PerplexityBot and Perplexity-User through robots.txt; again, publication is not evidence of consumer-side weighting. Google Search: documents ordinary Search eligibility and SEO foundations for AI Overviews and AI Mode, not a special llms.txt signal.
Key Takeaway: Separate crawler access, vendor self-publication, custom-agent capability, and proven citation effects. They are not interchangeable evidence.
How Do You Write a Clear, Useful llms.txt File?
A useful llms.txt is concise, accurate, maintainable, and explicit about which resources matter. This seven-step workflow follows the v2 proposal without promising platform adoption.
Step 1: Choose scope. Use /llms.txt for a site-wide map or a path-scoped file such as /docs/llms.txt when a section has its own audience and resources.
Step 2: Write the H1 and summary. Name the project, then explain what it is, who it serves, and what the linked resources cover.
Step 3: Add only necessary guidance. Include stable facts or instructions that help an agent interpret the links. Do not insert promotional claims that cannot be verified on the linked pages.
Step 4: Group links under descriptive H2s. Match the information architecture users understand: Docs, API, Guides, Services, Research, or another clear label.
Step 5: Curate canonical URLs. Prefer the pages that define the product, organization, evidence, policies, and subject expertise. Avoid dumping the whole sitemap.
Step 6: Add accurate optional notes and discovery links. Notes after each Markdown link can explain scope. Where useful, expose Markdown twins and the v2 alternate/describedby relationships.
Step 7: Publish, validate, and maintain. Return HTTP 200 with readable Markdown, check every URL, exclude sensitive material, and review the file whenever canonical resources change.
Key Takeaway: The useful work is scope, curation, accuracy, and maintenance, not claiming that an undocumented platform must honor the file.
The fastest creation handoff is the free llms.txt generator: generate a first draft, review every URL and description, then publish it through your site's normal deployment process.
llms.txt vs llms-full.txt: When to Publish Both
llms.txt is a compact map; llms-full.txt is a community convention for a combined full-text export. The v2 proposal emphasizes following links to LLM-friendly content and no longer defines special context-expansion tooling as part of the proposal.
A documentation-heavy product may still publish a combined file for agents or customers that explicitly want one-document ingestion. Most brand, agency, ecommerce, and editorial sites should begin with the compact map and clean Markdown versions of their most useful pages. Do not publish a large combined file merely because a model advertises a large context window: size limits and retrieval behavior vary by product, and no universal threshold guarantees ingestion.
If you publish both, link the full export clearly, keep it generated from public canonical content, and apply the same privacy, stale-link, and maintenance checks used for the main manifest.
How Long Should an llms.txt File Be?
Do not design llms.txt around a model's advertised maximum context window. Context limits, retrieval steps, and attention behavior differ across products and can change without notice. A concise map is easier to inspect, maintain, and use than a full-site dump.
Prioritize in this order: a stable project description; canonical product, service, policy, and evidence pages; task-specific documentation; then optional supporting material. Keep each description factual and distinct. If the file becomes difficult for a human editor to review, split it by path or move long-form content into linked Markdown twins rather than inventing a universal token budget.
Key Takeaway: Optimize for clarity and maintainability, not an unsupported token threshold. V2 path scope and Markdown twins provide cleaner ways to manage large sites.
What to Exclude From Your llms.txt (Security and Reputation Guardrails)
Auto-generators that concatenate an entire sitemap into llms.txt or llms-full.txt routinely pull content that should never be handed to an AI engine: exposed API endpoints, staging URLs, executive personal contact details, and unredacted client data all leak this way. A curated exclusion list is not optional for any production deployment.
| Content Vector | What Good Looks Like | Common Destructive Mistake |
|---|---|---|
| Product documentation | Clean markdown explaining capabilities, use cases, and public API surface | Accidental inclusion of internal API endpoints, admin routes, or leaked developer keys |
| Executive information | High-level biographies and public-facing thought leadership | Exposing direct cell numbers, personal emails, or home-city references pulled from bios |
| Sales narratives | Neutral, educational comparison copy with verifiable claims | Aggressive "buy now" language that AI engines filter as promotional spam |
| Customer evidence | Anonymised, statistically sound case-study outcomes | Leaking sensitive client usage data, contract terms, or logo permissions without prior redaction |
| Environments | Production URLs only, served with correct canonical tags | Staging subdomains, preview URLs, and internal QA paths concatenated into the file |
Bake the exclusion list into whatever script generates the file; do not rely on manual review of a 30,000-line output. A single explicit deny-list in the build step (staging subdomains, /admin, /internal, /api/private, user-generated comment feeds, any URL containing customer identifiers) prevents almost every real-world leak we have audited.
Key Takeaway: Treat llms.txt as a curated public library, not a database dump. An explicit exclusion list in the generator script (staging, admin, private API paths, PII, and unredacted client data) is mandatory for any production deployment.
Real-World llms.txt Examples (And What They Get Right)
Published files are useful examples of producer adoption, not proof that the publisher's AI products consume third-party llms.txt files. Review live files because their contents can change:
- Anthropic developer documentation. A documentation-oriented map produced by the company behind Claude.
- OpenAI developer documentation. A curated API-docs index with Markdown twins and a combined export.
- Perplexity developer documentation. A structured index for its API documentation.
- Smart Money Media. A site-wide map of services, guides, tools, research, and editorial resources.
The transferable lesson is curation: organize around the tasks and questions an agent may need to resolve, describe links accurately, and keep the file aligned with the canonical site. Do not infer consumer-product support from any company publishing the format for its own documentation.
Key Takeaway: Examples show how organizations produce the format. They do not establish which ranking or citation systems consume it.
How to Test Whether AI Engines Are Actually Reading Your llms.txt
Testing should separate technical validity, observed access, and downstream outcomes. None of these alone proves that llms.txt caused a citation.
1. HTTP and format validation
- Request the applicable root or path-scoped file and confirm HTTP 200.
- Confirm the response is readable Markdown, not an HTML error page or forced download.
- Validate the H1, links, optional notes, and path scope against the v2 proposal.
- Check every linked URL for a successful response, canonical accuracy, and accidental private or staging content.
- Use the free AI Crawler Check to review access rules separately from llms.txt syntax.
2. Access-log observation
Search server or CDN logs for requests to the file and record the user agent, timestamp, status, and path. Verify claimed bots against official documentation or published IP ranges where available. A crawler hit proves that a request reached the file. It does not prove parsing, weighting, ingestion, retrieval influence, or citation causality.
3. Outcome monitoring
Maintain a stable set of brand, product, and documentation questions. Record answer accuracy, cited URLs, and changes over time across relevant platforms. Treat any correlation cautiously because model updates, search indexes, page changes, and third-party coverage can all change the result. A controlled before/after test can generate a useful observation, but it still does not establish a universal platform rule.
Key Takeaway: Validate the file, observe access, and monitor outcomes as three separate evidence layers. Never turn a crawler request into a citation claim.
For the foundations this file cannot fix (crawlability, entity clarity, public evidence, and authority) use the secondary AI Visibility Audit.
Industry-Specific llms.txt Patterns
Section names should reflect the resources an agent or user would actually need, not a generic template.
SaaS and developer products: prioritize Docs, API, Security, Integrations, and stable troubleshooting resources. Consider path-scoped files for large documentation areas and Markdown twins for key references.
Agencies and professional services: prioritize Services, Methodology, Research, Case Studies, About, and editorial guides. Keep claims neutral and link to the evidence that supports them.
Ecommerce: prioritize durable category, policy, sizing, ingredient/material, care, shipping, and returns information rather than volatile individual inventory.
Media and publishing: prioritize subject hubs, original research, editorial standards, author pages, and evergreen explainers. Do not use the file as a replacement for a sitemap or feed.
Key Takeaway: A good structure mirrors stable information needs. V2 path scope is preferable to one oversized file when a site serves distinct audiences.
What Are the Most Common llms.txt Mistakes?
The most common failures are practical: an inaccurate file, a stale file, or a file that claims more than the evidence supports.
- Treating the proposal as a crawler-control standard. Access remains governed by robots.txt and other platform controls.
- Confusing self-publication with consumption. A vendor's own llms.txt demonstrates producer adoption, not a ranking or citation policy.
- Dumping the sitemap. A manifest should curate useful canonical resources rather than duplicate every URL.
- Missing or vague notes. Notes are optional in v2, but a precise description is more useful than an empty or promotional label.
- Stale, redirected, or noncanonical links. Review the file when URLs, products, policies, or documentation change.
- Ignoring path scope. Large sites can use the most specific applicable llms.txt instead of forcing unrelated resources into one root file.
- Publishing private or sensitive material. Never expose internal endpoints, credentials, personal data, client-confidential information, or staging URLs.
- Overclaiming validation. HTTP success and crawler requests do not prove parsing, weighting, or citation effects.
Key Takeaway: Accuracy, scope, curation, and evidence discipline matter more than rigid folklore about undocumented parsers.
How llms.txt Fits Inside a Full AEO and GEO Strategy
llms.txt is an optional technical workstream inside a broader visibility program. It can help package owned information for agents designed to use it, but it cannot create the underlying authority or eligibility those systems need.
- Crawlability and indexability. Public pages must be accessible through the controls each platform documents.
- Useful canonical content. The linked pages need clear definitions, complete answers, and maintained facts.
- Entity clarity. Names, categories, people, products, and structured data should agree across owned surfaces.
- Public evidence and earned authority. Independent reporting, research, and credible references corroborate owned claims. See the Reference Authority framework and media placements guide.
- Optional machine-readable curation. llms.txt and Markdown twins can make selected resources easier for compatible agents to navigate.
- Measurement. Track answer presence, accuracy, cited sources, and qualified buyer prompts without attributing a change to llms.txt unless the evidence supports it.
Use the LLM SEO guide for the full implementation context and the GEO guide for generative-search strategy. The primary next step on this page remains the llms.txt Generator. The secondary AI Visibility Audit checks the foundations llms.txt cannot fix: crawlability, entity clarity, public evidence, and authority.
Key Takeaway: Treat llms.txt as a low-cost curation layer. It is not a substitute for SEO, evidence, entity clarity, or earned authority, and it is not a proven ranking or citation signal.
Frequently Asked Questions
Common questions about llms.txt.
Sources & Further Reading
Use primary documentation to separate the proposal, crawler controls, and search eligibility. These sources document different things and should not be treated as proof of third-party llms.txt consumption:
This guide was drafted with AI assistance and edited, fact-checked, and approved by the Smart Money Media Team. Read our AI Use Policy.
Still have a question about this guide?
Ask Smart Money Media for cited answers pulled from our library, grounded in this article.
Create a clear, current llms.txt for your site.
Generate an editable v2 draft, review its canonical links and descriptions, then publish it through your normal site workflow.
Latest llms.txt Articles
Fresh insights and tactical deep-dives published in the llms.txt cluster.
How to Create a Wikipedia Page for a Business
Discover the strict editorial guidelines, conflict of interest rules, and step-by-step processes required to establish your company's digital encyclopedia presence.
How to Write a Pitch Email to a Journalist (+ Templates)
Securing top-tier press coverage requires mastering how to write a pitch email to a journalist. Learn the proven anatomy, templates, and strategic cadence.
How to Get a Google Knowledge Panel (Fast & Free)
Securing your entity in search changes everything. Learn how to get a google knowledge panel by mastering schema, third-party citations, and entity SEO.
Newsjacking: How to Hijack Trending News for Free PR
Master the strategic PR practice of capturing media attention during breaking stories. Learn how to safely insert your brand into trending narratives.
PR Measurement: The Definitive Blueprint for Brands
Stop relying on vanity metrics. Learn how modern PR measurement connects earned media and brand visibility directly to pipeline and business outcomes.
Digital PR Agency: What It Does and How to Pick One
Most brands remain invisible to search engines and AI assistants. A digital PR agency turns your best editorial placements into verifiable search authority.