Barely two years ago nobody had heard of it, and today more and more clients are asking whether they need an llms.txt file. It is a fair question, because two contradictory voices dominate the topic. One says this is the next big thing of the AI era, the other says it is a pointless technical exercise celebrated more by marketers than by search engines or AI systems.
The short answer: according to a well-documented picture backed by primary sources, llms.txt today is a real but narrow-scope tool. It does not improve Google rankings, as both a Google employee and Google's official guidance state. And there is currently no adequate evidence that it would, on its own, measurably increase a company's visibility in ChatGPT or Gemini answers: none of the major AI providers documents that it routinely reads an arbitrary website's llms.txt file before answering. Where it does actually work is in the world of developer and agent-based tools, for documentation-centric products and APIs. So the relevant question is not whether you need llms.txt, but whether your website would benefit at all from that narrow but real use case.
This article unpacks exactly that: what llms.txt really is, what we actually know about it with proof, what is only assumed, and which companies should bother with it at all.
Does your website need llms.txt in 2026?
Before we get into the details, here is the short answer by website type — it does not replace case-by-case assessment, but it is a good starting point.
What this ranking does not hide: the SEO effect is unproven, the AI Visibility effect is unproven, but agent readiness — that is, how easily an AI agent can work with a site's content — can have real value, primarily in the higher-priority categories of the table. Why these three concepts are not the same is explained in the next section.
SEO, AI Visibility and agent readiness: three different things
Most of the debate around llms.txt stems from mixing three actually separate goals into a single question: do you need this file. It is therefore worth separating them sharply.
SEO: making a company findable on Google and traditional search engines, among the classic organic results.
AI Visibility: making a company's content appear, and be interpreted correctly, in the answers of systems like ChatGPT, Gemini or Perplexity when a user asks something.
Agent readiness: making sure that a developer or business AI agent, deliberately pointed at the company's content by someone, quickly and accurately finds what it needs there.
These three goals require different tools, adapt to the expectations of different systems, and based on what we currently know, llms.txt is meaningfully connected primarily to the third.
A recent development indirectly supports this: in 2026, the Chrome developer team introduced a new, still experimental category in Lighthouse called "Agentic Browsing", which measures how easily an AI agent can work with a given page — this includes WebMCP integration, accessibility structure, layout stability and, yes, the presence of llms.txt. The category's documentation describes being built for machine interaction: according to the llms.txt audit page, without the file agents may spend more time mapping the site's structure. The documentation makes no mention of search ranking or appearing in AI answers, and it labels the category as an experimental feature built on proposed standards. This fits well with the distinction we also use: how well a site performs in search, how much it appears in AI answers, and how well it can be processed by an agent — with llms.txt primarily connected to the latter.
Anyone who truly wants to prepare their website for the age of both machines and humans has to deal with all three layers, according to their own weight, rather than focusing on a single technical file.
What llms.txt really is, and who came up with it
llms.txt is a simple Markdown file placed in the root of a website, typically at example.com/llms.txt. The file starts with a mandatory H1 heading, optionally continues with a short summary blockquote, and then lists the website's most important content under H2 subheadings, with a short description for each link. The goal is for an AI system or AI agent to quickly understand what the site is about and where to find the most important information, without having to map the entire site structure.
The concept was proposed by Jeremy Howard in September 2024, with the idea that large language models process content differently than a traditional search crawler, so it would be worth providing a separate, clean format specifically intended for them. So llms.txt is essentially a recommendation: it tells a machine what is worth reading, not what it is allowed to access.
How it differs from robots.txt and the sitemap
The three files have three different roles, and the common misunderstanding stems precisely from lumping them together.
robots.txt regulates access: it tells a crawler which paths it may crawl and which it may not. sitemap.xml is a complete technical inventory of all relevant URLs of the site, optimised for machine processing, without editorial selection. llms.txt, by contrast, is an editorial selection: it neither grants nor denies access, and it is not a complete inventory either, but a narrow, hand-picked list of what matters most.
robots.txt opens or closes the gate, the sitemap maps the house, and llms.txt gives a short guided tour of the most important rooms.
Official standard, or just a community convention
This is one of the most misunderstood points, so it is worth quoting the official source, the llmstxt.org website: in the project's own words it is still "a proposal to standardise", that is, a standardisation proposal, not an adopted, official standard. There is no standards body behind it, no binding specification that search engines or AI companies would have to comply with. The format is maintained by the llmstxt.org project and a few players involved in the topic, based on community consensus, not on any official obligation.
This does not mean it is worthless. Many well-functioning web practices started exactly like this, as community conventions, and over time either became widely adopted or remained marginal. llms.txt is currently in this in-between state, and that is exactly what justifies a cautious, fact-based approach instead of anyone uncritically treating it as mandatory.
What changed in llms.txt v2
The second version of the specification was released in August 2026, which, in Jeremy Howard's own words, was updated based on two years of adoption experience. v2 modifies the original 2024 version in four points.
First: it introduces standard link relations that, in an HTML `<link>` element or an HTTP header, help an AI agent find the Markdown version of a page or its associated llms.txt file.
Second: it makes the path of Markdown versions more flexible — the original required an .md extension appended to the full URL, while v2 also allows the extension to replace the original one.
Third: it clarifies the behaviour of an llms.txt file placed in a subfolder: the most specific file covering its own path always takes precedence. This matters for those who only control a subdirectory, for example project documentation running on GitHub Pages.
Fourth: it drops the earlier processing logic built on a specific helper tool (llms_txt2ctx) and phrases things more generally: AI agents search through llms.txt and then follow the relevant links. This also shows where the format is heading: less tied to a specific tool, becoming more of a general convention.
What we know for sure in 2026, and what we don't
Since a great many misleading claims circulate around the topic, it is worth separating what can be backed by documents or primary sources, what is promising but still early practice, and what currently lacks adequate evidence.
Proven, documented facts. Google officially and openly rejects the idea that llms.txt plays any role in search ranking. John Mueller, a member of Google's Search team, replying to a Bluesky post on 20 January 2026 to the question of whether the llms.txt file found on Google's own developer pages meant the format had been adopted, wrote: "I'm tempted to say something snarky since this has come up so often, but to be direct, no." In other words, an explicit no. Google Search Central's official guidance on generative AI features also states this (last updated: 10 July 2026): Google Search does not use such files, so creating them neither harms nor helps visibility and ranking on Google. At the same time, the guidance also says it is perfectly fine to maintain an llms.txt file for other services or systems. As early as June 2025, Mueller wrote in a Bluesky post that currently no AI system uses llms.txt. In addition, an Ahrefs analysis published in June 2026, which examined the May 2026 traffic and bot traffic data of 137,210 domains, found that 28 percent of the sites studied publish an llms.txt file, but 97 percent of these did not receive a single visit in the month examined. Where the file did receive requests, 96 percent of the traffic came from bots, and all AI-type bots combined accounted for only 19.5 percent of requests, with AI bots retrieving content specifically for answering making up a mere 1.1 percent of all requests. The research also recorded that no AI bot attempted to look for an llms.txt file on sites where it did not exist at all, meaning that proactive, automatic searching for the file is not typical behaviour today.
Important context is that the Ahrefs sample is not a random cross-section: it is skewed towards technically more active, more SEO-conscious domains, which may push the adoption rate upwards: Ahrefs itself considers the 28 percent figure an upper bound. This is well illustrated by Rankability's measurement of 18 September 2026, which examined the root files of the most popular domains based on the Tranco list: 9.3 percent of the top 1,000 sites and 8.3 percent of the top 10,000 sites serve a valid llms.txt or llms-full.txt file — so on a broader sample that is not SEO-conscious to begin with, the rate is much lower than in the Ahrefs sample. Broken down by industry, travel (16.0 percent) and financial services (15.4 percent) lead, while the technology sector stands at 6.4 percent. Rankability's earlier June measurement found 8.7 percent for the top 1,000 sites, but the sampling and validation methods have changed since, so the two figures cannot be compared directly.
Emerging practice, promising direction. A growing number of documentation-generating platforms, typically sites serving developer tools, automatically generate an llms.txt file for their own documentation, and a separate tool category is also emerging that makes llms.txt-like, clean documentation available to coding AI agents, typically through MCP-based integrations. This practice is real and growing, but tool-dependent: not every development environment reads it in automatically; it requires deliberate configuration or a separate add-on.
What currently lacks adequate evidence. There is no documented data backed by primary sources showing that the presence of llms.txt on its own would increase how often ChatGPT, Gemini, Perplexity or Claude cite a company in their answers. There is no official statement from OpenAI, Google or Anthropic saying that their own systems routinely read an arbitrary third party's llms.txt file before answering. What is documented: all three companies publish an llms.txt file for their own developer documentation (OpenAI, Anthropic, Google) — this shows that they offer the format to the developers who use them, not that their search or answering systems automatically read other sites' llms.txt files.
AI training, AI search and AI agents: why the three are not the same
There is also a technical reason behind the threefold division above: AI systems encounter a website's content in three very different ways.
The first is model training, which takes place on huge, pre-collected datasets, long before a user's question is asked — a file placed in a website's root has practically no effect on this.
The second is search and retrieval running at the moment of answering, when ChatGPT, Gemini or Perplexity pulls in web results in real time. Here, technical accessibility, well-structured HTML, a clear content structure and technical SEO lay the foundation, and here, according to the data cited above, the role of llms.txt is negligible today.
The third is the world of developer and business AI agents: tools that a person launches for a specific task, explicitly, along a given URL or integration. This is the layer where llms.txt can actually be useful, because this is not about automatic, mass crawling, but about targeted queries initiated by a human.
These three layers respond to an llms.txt file to different degrees, and anyone who does not separate them can easily draw the wrong conclusion about the format's usefulness in either direction.
Who should bother with it: the W5labs decision matrix
Instead of saying yes or no in general, we examine a given website along two criteria. We use this framework under the name W5labs llms.txt Priority Matrix.
The first axis is how documentation- or API-like the content is: is there structured, repeatedly used knowledge that developers, integrators or AI agents regularly look up. The second axis is how much the business model relies on direct, repeated content consumption by AI agents, as opposed to the vast majority of visitors being humans arriving via traditional search or direct visits.
The matrix breaks down the ranking seen at the beginning of the article along two axes: API documentation and developer tools typically fall into the top-right quadrant, while traditional company presentation and landing pages fall into the bottom-left. The bottom-right quadrant deserves special attention: there the question is not the file, but whether the site can be processed well by an agent at all. The framework is a decision-support tool, not a measurement result, and it does not replace case-by-case assessment, but it is a good starting point for a company to see realistically how much is at stake in its own situation.
What a good llms.txt file should contain
A good llms.txt is a short, clean Markdown file: a concise introduction to the company or product, followed by links to the most important content, each with a one-line description.
Bad practice is when a company simply copies its entire navigation menu without descriptions, producing a long link list without context, from which an AI agent can no more tell what matters than if there were no file at all.
Good practice is when the file is a deliberate selection of a few truly relevant pages, with concise and specific descriptions that on their own tell a reader what they will find there.
The difference is not in the length of the file, but in the fact that in the second case a machine also knows exactly which link is worth clicking and why.
Typical mistakes
The most common mistake is when a company copies its marketing copy into llms.txt instead of factual, concise descriptions. The second is outdated content: the file is created and then not updated for years, while the site's structure changes and some of the links become invalid. The third is excessive expectations: people expect short-term, measurable SEO or AI visibility improvements from it, then are disappointed when these fail to materialise, even though — as shown above — the file on its own is not suited for that. The fourth is when the company puts the emphasis on llms.txt while neglecting the maintenance of robots.txt and sitemap.xml, even though these carry considerably more weight in terms of technical fundamentals.
llms.txt is not a GEO strategy, just one technical element of it: the W5labs approach
It is important to see clearly: llms.txt on its own is not an AI Visibility strategy, nor does it replace one. A file in the root directory does not solve issues such as content depth, the correct use of structured data, the site's technical indexability or the company's external credibility signals. At most, it is a complementary, occasionally useful layer within a much broader piece of work.
That is why, for us, the question of llms.txt does not start as a standalone technical project, nor with technical SEO. In our Strategy → Creative → Development → Support model, the Strategy phase assesses the current situation: the company's business goal, its SEO and AI Visibility/GEO status, and sets the priorities based on this. For this we use our own measurement framework, GVS (Growth Visibility Score) — it helps quantify the current situation and identify the most important gaps and opportunities, and based on these it can be determined how much priority llms.txt deserves at the given company.
The tasks that emerge this way are then passed on to the further phases of the model: strengthening indexability, structured data, content depth or external credibility signals is the joint work of Creative and Development, the technical implementation of llms.txt is specifically a Development task, while its verification and maintenance are part of ongoing Support. Whether it is needed at all, and with what priority, however, remains a Strategy question.
We like to say that not every new thing needs to be introduced immediately just because it can be.
What we are testing now, and why we are not saying more than this
W5labs is currently preparing its own longer-term observation, in which we will examine, in a few controlled cases, how a carefully built llms.txt file behaves on real traffic data. In the observation we track how often and which bots request the llms.txt file and the pages it references, how this traffic is distributed between answer-oriented AI bots, training crawlers and developer agents, and whether in a targeted agent task an agent reaches the information it is looking for faster or more accurately with the file than without it. In parallel, in our GVS measurement we monitor the AI mentions and Share of AI Voice of the sites studied, compared with control sites that have no llms.txt. This work is coming together now and has no results yet, so at this point we are deliberately not claiming anything with reference to our own measurements. As soon as we have enough data to report on it responsibly, we will share it in a separate article.
llms.txt is a real but, for now, limited-scope tool. It is not an official standard, but a continuously evolving community proposal whose second version was released in August 2026. It has no documented role in classic SEO, as Google has also confirmed. In consumer AI search and answering, its actually measurable impact is minimal according to current data. Where it does bring real value is in the targeted, deliberate use of developer tools and AI agents, primarily for documentation-centric products and APIs. This means that most companies do not need to jump on it right now, but for a specific group, depending on their content and business model, the investment is genuinely worth it.



