As of this article’s date, no generative engine has publicly stated that it consumes llms.txt as a discovery or citation channel, and the specification itself does not ask for that. Publishing a file containing accurate public content is inexpensive; keep it up to date. Relying on it as a mechanism for appearing in answers means betting on a channel nobody has confirmed exists.

I write from murmur.marketing, a GEO (Generative Engine Optimization) practice in Brazil, operated by Mateus Gomes. This is a dated, falsifiable position: the day OpenAI, Anthropic, Perplexity or Google publishes a statement that it reads the file, this page will be wrong and I will rewrite it.

Summary

  • Jeremy Howard proposed the specification on 2024-09-03. It explicitly states that it does not recommend how the file should be processed: that decision belongs to the consumer.
  • Its stated purpose is context at inference time, not crawling. Those are different stages, and confusing them is where the folklore begins.
  • None of the four engines I measure — ChatGPT, Claude, Perplexity and Google AI Overview — has published documentation about consuming the file.
  • In a server log from a client’s subdomain, the legitimate crawler that read most requested robots.txt 15 times, the two sitemaps 34 times combined, and article pages 24 times. Those are the channels visible in the log.
  • The three established discovery paths remain internal links from known pages, a sitemap declared in robots.txt, and active submission to an index.
  • The practical recommendation: publish it, but do not rely on it. It takes ten minutes and barely registers in a spreadsheet of the work required.

What the specification actually says: less than the market repeats

Jeremy Howard, also known for fast.ai and Answer.AI, proposed the file on 2024-09-03. The proposal is short, and the problem it describes is specific. In translation of the passage quoted in the Portuguese article: “Language models increasingly depend on website information, but face a critical limitation: context windows are too small to handle most websites in full.”

Read that sentence again: it determines the issue. The stated problem is the context window, how much text fits into the model when it is already composing an answer. It is not discovery, crawling or indexing. The specification proposes curated Markdown to help the model use a site at inference time.

There is a second sentence almost never quoted alongside it. Again, translating the passage cited in the Portuguese article: “This proposal does not include any particular recommendation for how to process the llms.txt file, since that will depend on the application.” The specification deliberately declines to define consumer behavior. A standard that does not specify how it is consumed cannot, by construction, promise that someone consumes it in a particular way.

None of this criticizes the author. It is an honest proposal addressing a real problem. What happened afterward was its repackaging as a GEO checklist item. The distance between “helps the model read my site if it is already reading it” and “makes AI cite me” is the distance between two different stages of a two-stage system.

Two stages, and why confusing them is the whole error

A generative engine with search does two things in sequence: it retrieves documents, then generates text from them. The mechanics are explained in What GEO is, and the terminology I use to separate the stages is explicitly labeled as my own in Retrieval and generation in AI search.

StageThe question it answersWhat affects itWhere llms.txt fits
DiscoveryDoes the crawler reach the page?Internal links · sitemap.xml in robots.txt · active submissionNo evidence that it plays a role
RetrievalIs the document selected from the candidates?Page specificity and matching the questionNo evidence that it plays a role
GenerationDoes the brand survive the writing of the answer?An answer at the start, a named entity, numbers with denominatorsThis is where the specification places itself
Entity recognitionDoes the engine identify you as a distinct entity?Third-party sites and sourcesOutside the file’s scope

The file presents itself in the third row. Most market advice sells it in the first. That is why the honest answer to the title is neither “yes” nor “no,” but “you are asking about the wrong stage.”

What the logs show actually being requested

Here I replace speculation with server records, the least glamorous and hardest-to-dispute evidence in this category.

In a 24-hour window on a client’s subdomain, I cross-checked 2,256 nginx access-log lines by user-agent and real IP address. Raw user-agent counts showed ClaudeBot 88, OpenAI 10 and Perplexity 6. IP checks left ClaudeBot 79, OpenAI 3 and Perplexity 0. The rest was credential scanning disguised as crawler traffic, requesting /.env and /.git/config with seven forged user-agents sent in the same second.

Of the 79 legitimate ClaudeBot hits, the requests included 24 article hits covering 20 distinct slugs, 17 for the content sitemap, 17 for the sitemap index and 15 for robots.txt. This is a reading pattern, not a scan. The infrastructure files it repeatedly consults are precisely the two described in public crawling documentation.

⚠️ The caveat this observation requires is mine to state: I did not run a controlled test serving llms.txt and measuring requests for it. The log shows what was requested, not proof that the file would not have been requested if it existed. This indicates where crawler traffic is concentrated today; it is not an experiment. I treat it as an indication and ask readers to do the same.

The broader scale comes from an external source and is substantial. In Vercel’s network logs, AI crawlers combined — GPTBot, Claude, AppleBot and PerplexityBot — made roughly 1.3 billion requests in one month, around 28% of Googlebot’s volume on the same network. Real AI crawler traffic exists. It simply is not taking the path the folklore describes.

Three discovery channels that hold up

If the aim is to be reached, focus on these three, in order of cost-effectiveness:

  1. An internal link from a page the engine already knows. The cheapest and most overlooked path. An orphan page does not discover itself.
  2. A sitemap.xml declared in the root domain’s robots.txt. The easily missed detail: crawlers read robots.txt at the root, and only there. Nobody fetches one published in a subdirectory. I nearly made this mistake on my own site: the domain’s robots.txt did not exist and returned 404. A section-level file cannot substitute for it.
  3. Active submission to an index. ChatGPT with search is grounded in Bing’s index, making Bing Webmaster Tools an opportunity almost nobody in this market uses.

A technical reminder that invalidates more sites than any new file could fix: AI crawlers do not execute JavaScript. Content existing only after hydration does not exist for them. There is no second chance. If the served HTML is empty, the world’s best-written llms.txt has nothing useful to point to.

Why I publish the file anyway

I do publish it. The reason is asymmetric cost, not conviction.

Writing llms.txt takes ten minutes when the public content is reviewed and maintenance is planned. If an engine starts consuming it, check its consumption rules and whether the file is current. If none does, the cost was ten minutes. That is an inexpensive good-faith step, distinct from treating it as a channel.

The line I do not cross is charging for it as a deliverable. A vendor listing llms.txt as a scope item is selling something whose effect nobody has demonstrated. That is a legitimate buyer question, not nitpicking. The standard I apply to my own work is explained in How to tell whether a GEO agency has results.

The state of the debate, and my position within it

Disclosure of interest: murmur delivers GEO as SWAS, combining proprietary software for measuring and tracking SOV with specialists in strategy and execution. The August 6, 2026 campaign covered the Brazilian market with 100 Portuguese questions across four engines, 800 captures, all with screenshots. The classifier recorded citation_kind = ausente in 800 of 800; outside the ten questions already naming the brand, the textual count was 0 of 712. This is a baseline before the current guide library, not an assessment of today’s strategy or an efficacy test of llms.txt.

Across the category, my deterministic rule — applied to a declared list of 29 names, across 100 (question, engine) pairs on 2026-08-06 — found 51 of 100 pairs with no winner at all. The most frequently cited name, Conversion, wins 19 of 100 pairs and does not reach a majority in the other 81; GeoStack wins 15 and Brasil GEO 14. The top three are separated by less than 1.3 binomial standard errors, so this is not a stable ranking. These figures do not test llms.txt. Assess the file through documentation and observed consumption, not the ranking of whoever recommends it.

External skepticism is healthy and specific. The critical survey arXiv 2607.14035 (Olivier Martinez, 2026-07-15) reviews 45 studies published between November 2023 and July 2026, concluding that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect. It classifies the literature into five evidence levels and places the category’s founding paper at level C. Christopher Penn approaches the criticism through instrumentation: much of the measurement advice sold here comes without a method. llms.txt exemplifies both problems: a technique without a demonstrated effect, recommended by people who do not publish how they measured it.

In the historical record for Fly Vet and Fly Med, owned by the same operator as murmur, Fly Med received its first ChatGPT citation of its /geo/ directory at T+5 days after publication: two pages in one query, with ?utm_source=chatgpt.com in the URLs. Fly Vet recorded zero in the same probe. Neither had llms.txt. The contrast shows that the observed citation did not require the file; it does not isolate the cause of the difference between the companies.

Frequently asked questions

Does the llms.txt specification promise that an engine will read the file?

No, and it is explicit about that. The text published at llmstxt.org says the proposal makes no recommendation about how the file should be processed, because that depends on the consuming application. It proposes a format, not a behavioral contract. Turning that into a citation promise adds a guarantee the original document deliberately declines to make. That is intentional, not an omission.

What is the practical difference between llms.txt and robots.txt?

They belong to different stages and have different evidential status. Crawler operators document robots.txt behavior; it is repeatedly requested at the domain root and is where the sitemap is declared. In a 24-hour log from a client’s subdomain, the legitimate crawler requested it fifteen times. The purpose of llms.txt is to help a model already reading the site fit the material into its context window, and no operator has published consumption documentation. One is crawling infrastructure verifiable in logs; the other proposes a format for a different stage.

If I already published llms.txt, should I remove it?

If the public content is reviewed and there is a maintenance plan, there is no reason to remove the file solely because it is llms.txt. Reassess it if it contains outdated, sensitive or non-public information; assign an owner and review links and content regularly. Adjust your expectations and planning instead: move it out of the acquisition-channel column and into technical good faith. The problem was never that the file exists. It was letting the file displace work with demonstrated effects, such as serving HTML without JavaScript dependencies and declaring the sitemap at the root.

What would change my position on this?

Public documentation from an engine operator stating that it consumes the file, or a request appearing in server logs with the operator’s verified IP. I would accept the first immediately; the second is something I could measure myself. A larger volume of advice would not change my position: how many people recommend a practice is not evidence that it works. This entire category still operates in a field where 51 of 100 (question, engine) pairs have no dominant brand.

Can publishing llms.txt cause harm?

The impact depends on what public content the file contains and how it is maintained: review claims, data, URLs and permissions before publishing, and keep it current. This analysis did not measure a specific effect of the format. The allocation risk remains: the file is easy to create, feels like progress, and competes for attention with tedious work known to matter — serving complete HTML without hydration, declaring the sitemap in the root robots.txt, and submitting to Bing’s index, which grounds ChatGPT search. If llms.txt is the week’s completed task while those three remain unresolved, it was expensive despite costing nothing.

Who wrote this, and disclosure of interest

Mateus Gomes, operator of murmur.marketing, a Brazilian SWAS GEO operation. This commercial interest is disclosed; the technical recommendation separates preparing a file from proving its effect. llms.txt may be part of technical content organization without being presented as a proven mechanism for increasing citations. SOV-growth targets and guarantees depend on the contract model, scope and conditions.

Conclusion

llms.txt is an honest proposal about context windows that the market has converted into a promise of citations. The specification makes no such promise, explicitly declines to define how the file should be processed, and no engine operator has published consumption documentation. Server logs repeatedly show requests for robots.txt and sitemaps. Consider publishing the file if its public content has been reviewed and a maintenance plan is in place; creation can take about ten minutes. Without those conditions, focus your effort on the three established discovery channels, HTML served without depending on JavaScript, and entity recognition, which no file on your own domain can resolve. To discuss your site’s situation, contact Mateus Gomes on LinkedIn.

See also