Technology News

We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read

Everyone has an opinion on llms.txt, but when it comes to actual evidence we have only single-site logs or the odd small-scale experiment.

Using Ahrefs Web Analytics and Bot Analytics, we analyzed the server logs and live traffic of 137K domains, plus the user agents hitting all of them.

Here’s what we found.

Top findings

  • 28% of the 137K domains using Ahrefs Web Analytics publish an llms.txt file.
  • 97% of those files received zero traffic in May 2026. Nothing fetched them at all.
  • 96% of the requests that did reach llms.txt files came from bots.
  • 19.5% of fetches came from named AI tools (of the 3% of files that weren’t ignored). GPTBot is top and Claude-Code is second, ahead of every AI search and assistant bot.
  • 12% of fetches come from the industry studying itself: GEO/AEO tools, llms.txt checker tools, and researchers.
  • Zero requests came from AI bots for llms.txt files that don’t exist. They never go looking.
  • The Chrome Lighthouse llms.txt audit produced roughly 1 in 1,000 fetches.

In late May 2026, Google took both sides of the llms.txt argument in under a week.

Its new guide on optimizing for generative AI features told site owners, in a section literally titled “mythbusting”, that machine-readable files like llms.txt aren’t needed to appear in generative AI search.

A text excerpt from "Mythbusting generative AI search: what you don't need to do." Highlighted text states you don't need special files or markup for generative AI search.

 

Days later, the Chrome team shipped an llms.txt check inside Lighthouse’s experimental Agentic Browsing audits, with documentation explaining that without the file, agents may spend more time crawling a site to understand its structure

Agentic browsing audits" section.” data-src=”https://technobabble.com.au/wp-content/uploads/we-analyzed-137k-sites-97-of-llms-txt-files-never-get-read.png” srcset=”https://technobabble.com.au/wp-content/uploads/we-analyzed-137k-sites-97-of-llms-txt-files-never-get-read.png 2048w, https://technobabble.com.au/wp-content/uploads/we-analyzed-137k-sites-97-of-llms-txt-files-never-get-read-7.png 607w, https://technobabble.com.au/wp-content/uploads/we-analyzed-137k-sites-97-of-llms-txt-files-never-get-read-8.png 768w, https://ahrefs.com/blog/wp-content/uploads/2026/06/a-webpage-titled-llms-txt-on-chrome-for-develope-1536×1076.png 1536w” sizes=”(max-width: 2048px) 100vw, 2048px”>

 

When Lily Ray pressed Google’s John Mueller on the contradiction, he explained that llms.txt is “not done for search.” It’s a “temporary crutch, perhaps to save some tokens” for AI coding tools parsing developer documentation—not something non-developer sites need to worry about.

He also stated that site owners who check their logs will find very little AI agent traffic.

A screenshot of a Twitter thread from John Mueller. The highlighted text says, "even with more agentic traffic in the future (and if you check your logs, you’re not getting a lot of that at the moment)."A screenshot of a Twitter thread from John Mueller. The highlighted text says, "even with more agentic traffic in the future (and if you check your logs, you’re not getting a lot of that at the moment)."

 

This is something we decided to test.

 

What llms.txt is (and what it isn’t)

Before we go any further, let’s clear up what llms.txt actually is. Llms.txt is a single index file, written in markdown, placed at a site’s root. Proposed by Jeremy Howard, co-founder of Answer.AI and fast.ai, in 2024, it summarizes what a site is and links its most important content. The idea being that LLMs and agents can use this information to orient themselves without crawling everything. The “AI visibility” framing around llms.txt came later on, attached by the SEO industry as adoption spread on the speculation that AI platforms would reward the file. Two things it is often confused with, and isn’t.

  • It is not the practice of publishing markdown copies of your web pages, a separate tactic with its own problems.
  • And despite the filename, it is not a robots.txt-style directive: it controls nothing and blocks nothing.

This study measures the index file, and only the index file. 

Methodology 

Our study focuses on all 137,210 domains in Ahrefs Web Analytics that received traffic in May 2026.

We checked each domain root for an llms.txt returning HTTP 200, then used Ahrefs Bot Analytics to examine every request to /llms.txt paths across the population, split by HTTP response (200 vs 404) and classified by channel and individual user agent.

To rule out soft 404s and phantom files, we also confirmed each file was actual Markdown rather than HTML, and screened titles and content for error signals like “404” or “Page not found”

It’s important to note:

  • Ahrefs Web Analytics customers skew more technical and SEO-aware than the web at large, so treat the 28% adoption figure as an upper bound.
  • We did not explicitly study whether a file was well-formed against the llms.txt specification.

Chris Long, Founder of Nectiv, has also stated that, even if llms.txt doesn’t help you in Google search, the file has utility if your customers “are using Claude Code to source recommendations”

LinkedIn post by Chris Long about LLMs.txt and its relevance to SEO beyond Google Search, with highlighted text.LinkedIn post by Chris Long about LLMs.txt and its relevance to SEO beyond Google Search, with highlighted text.

Our Bot Analytics data supports both ideas.

We see llms.txt files being fetched far less by the search and AI bots that are seemingly responsible for visibility, and far more by the agentic tools that seek out structured information and/or act on a user’s behalf.

Bar chart showing the share of verified AI bot requests from various agents, totaling 10.5%. "statespace-indexer" leads with 3.52%.Bar chart showing the share of verified AI bot requests from various agents, totaling 10.5%. "statespace-indexer" leads with 3.52%.

*statespace-indexer: operator identified as Statespace (agentic infrastructure), IP ranges unconfirmed.

Aside from statespace-indexer and GPTBot, Claude-Code (Anthropic’s coding agent), out-fetched every AI retrieval bot, every AI assistant, and every AI training crawler.

Training crawlers are the second-largest AI category at 5.3%

Llms.txt files feed training corpora more than they feed AI search retrieval.

In fact, AI training crawlers fetch llms.txt nearly 5X more than AI retrieval bots.

Bar chart showing 5.3% of AI bot requests come from AI training crawlers. GPTBot is 4.51%, ClaudeBot 0.8%.Bar chart showing 5.3% of AI bot requests come from AI training crawlers. GPTBot is 4.51%, ClaudeBot 0.8%.

So if llms.txt were to in any way impact your brand’s AI visibility, it would likely be upstream—not at the point of retrieval.

Of all training crawlers, GPTBot is far and away the biggest fetcher of llms.txt.

You won’t find a Gemini crawler in this list, because it doesn’t exist.

Google trains and grounds Gemini on content fetched by regular Googlebot, and Google-Extended, the opt-out publishers use, is a robots.txt token rather than a crawler with its own user agent.

Googlebot did fetch llms.txt files ~900 times in May, but Googlebot routinely fetches any URL it discovers on a site as part of normal search indexing, so those fetches don’t indicate special interest in llms.txt—it’s crawling the file the same way it crawls a sitemap or any other page.

Whether any of that content then feeds Gemini is invisible to us.

AI retrieval bots barely register, with 1.1% of total requests

According to our data, AI retrieval bots account for just 1.1% of AI bot requests.

Even when taken together with AI assistants and AI training crawlers, these bots still count for only 8.9% of requests (1.6% less than AI agents).

OAI-SearchBot, PerplexityBot, and Claude’s search crawler combined made only a couple of hundred fetches across thousands of sites.

Bar chart showing that 1.1% of AI bot requests come from AI retrieval bots. OAI-SearchBot leads with 0.74%.Bar chart showing that 1.1% of AI bot requests come from AI retrieval bots. OAI-SearchBot leads with 0.74%.

If you are planning on generating an llms.txt in hopes of boosting your AI citations, you may want to think again.

12% of requests come from tools studying llms.txt, not consuming it 

A whole ecosystem has formed around auditing, scoring, validating, and studying the llms.txt standard, before we’ve even established whether any major AI platform actually reads it.

Three categories account for 12% of all requests combined.

Pie chart showing 12% of requests study the llms.txt standard. Research bots: 2.7%, llms.txt discoverability: 3.6%, GEO/AEO tools: 5.8%.Pie chart showing 12% of requests study the llms.txt standard. Research bots: 2.7%, llms.txt discoverability: 3.6%, GEO/AEO tools: 5.8%.

 

GEO/AEO tools send 5.8% of requests

Commercial tools scan websites and score their readiness for AI search and agent discovery, with llms.txt presence as one of many signals.

The most active, CairrotReadinessBot, belongs to Cairrot, a WordPress-focused AEO platform launched in late 2025.

Then you have the mainstream website builders like Framer, Lovable, and Wix all baking AI-readiness checks into their products.

Lms.txt adoption has become a platform default before it’s even become a webmaster decision.

llms.txt discoverability bots cover 3.6% of requests

There’s an ecosystem of tools that catalog the llms.txt files that almost nobody else reads.

Dedicated scanners, validators, and directories built solely for llms.txt files send more requests than AI retrieval bots and AI assistants.

Research bots send 2.7% of requests

The largest single research crawler in the dataset identifies itself as prompt-injection-survey/1.0.

Someone is systematically studying llms.txt as a prompt injection opportunity that AI agents are designed to ingest and trust.

The security implications of agents trusting llms.txt files at scale have barely been discussed, and yet potential bad actors are already on the case.

Zero AI bots “go looking” for llms.txt files that don’t exist 

AI tools never go looking for llms.txt files that aren’t there, so publishing one does not put you on any AI radar.

We analyzed every request to /llms.txt paths that returned a 404 and found the cleanest split we’ve seen in bot data: where on the one hand valid files drew 96% bot traffic, missing files drew 98% human traffic, and the AI bot share of those 404s was zero.

The people probing for absent llms.txt files are humans typing the URL into a browser, presumably SEOs checking on competitors.

This kills the assumption that AI systems actively hunt for llms.txt files, and that a site without one is missing a knock at the door.

AI tools fetch llms.txt when a link, an index, or a user instruction tells them it exists.

How to check your own llms.txt bot traffic

If you want to see which bots are actually hitting your llms.txt file, head to Ahrefs Bot Analytics and add a filter for Page URL → Contains → llms.txt, then hit Apply.

studying llmstxt fetches in Ahrefs bot analyticsstudying llmstxt fetches in Ahrefs bot analytics

This narrows everything down to requests hitting your llms.txt file (or any pages with “llms.txt” in the URL, like blog posts about it).

We don’t have an llms.txt file on the Ahrefs site but we are getting some bots hitting that page, as indicated by the 404 status.

From there, you can check:

  • Visits over time. Toggle between By bot and By category to see whether traffic is climbing, flat, or spiking. 
  • The Bots table. See which exact bots are fetching the file.
  • Last status in Crawled pages. Check the status code. A 404 on /llms.txt means bots are asking for a file that isn’t there.

That last point is the useful gut-check. Plenty of sites get bot requests for an llms.txt they never published. The traffic is real; the file isn’t.

You can also use the AI bots filter at top of the page to strip out other crawlers and see only the LLM-related ones.

And, remember, a bot requesting your llms.txt isn’t proof anything read or acted on it. It only tells you the file was fetched.

AI search, there are more reliable ways to improve your visibility than this file.

But if you’re still toying with the idea of generating llms.txt, here are the steps you should take:

  1. Check your own logs before investing further. A 97% chance of zero readership is the base rate.
  2. Get a website-building platform to do it for you. Wix already generates these files, and Framer and Lovable are scanning for them. Within a year, having an llms.txt may be as much a CMS default as having a sitemap. If the payoff is uncertain, it makes sense to keep the effort minimal.
  3. Route agents to it. Link the file from your HTML, reference it in your docs, or mention it anywhere agents receive instructions about your site. Agents fetch llms.txt when directed, not speculatively.
  4. Offset the prompt injection risk by treating llms.txt like code. Version-control it, restrict who can edit it, set an alert for unauthorized changes, keep the content to plain links and descriptions (nothing instruction-shaped), only link to resources you control, and review anything a platform auto-generates on your behalf.

This study answers how many sites publish llms.txt, and who reads it.

But there are a couple of other questions worthy of further research that were beyond the scope of this study:

  1. Do agents fetch developer-docs more often? Is Claude-Code’s llms.txt interest concentrated on documentation paths like /docs/ and /api/, as Mueller’s framing predicts?
  2. Do bots actually act on what they read? When an AI agent fetches llms.txt, does it then fetch the resources the file links to? SEO consultant David McSweeney, Founder of Queryburst, is already running an experiment along these lines: he’s serving AI user agents a compressed, agent-friendly summary of his test sites, complete with instructions for requesting deeper content, and tracking whether any agent actually follows through. His results are worth following.

Mueller called llms.txt a temporary crutch.

But that crutch seems to already have its own supply chain: platforms generating llms.txt files, an industry auditing them, and security researchers studying them, all before the “readers” actually showed up.

Either we’re watching the early scaffolding of a real standard, or we’re watching the SEO industry prove it can productize anything. Our money is on a bit of both.

 

June 2026 Patch Tuesday: Microsoft Patches 206 Vulnerabilities Including Three Publicly Disclosed Zero-Days
Agentic Marketing: What’s the Big Deal and How to Get Started

Related Articles

Why does the WeWork guy get to fail up?

Why does the WeWork guy get to fail up?
Housing in the United States has a problem. And Adam Neumann, the charismatic founder known for successfully rebranding shared office space as WeWork, and unsuccessfully running it, thinks he has…

Design Studio & Agency Website Inspiration

design-studio-agency-website-inspiration
If you’re taking on the huge task of designing a studio or agency’s website, it’s important that you get everything just right. Professionalism and ease of navigation needs to balanced…