llms.txt: 162 Requests in 53 Days, Not One From an AI Crawler

This site has had an llms.txt file since 2026-08-29. I've written about it twice already: once about the first version being robots.txt with a different filename, and once about AEO and GEO making the same mistake at industry scale. Both pieces were about what the file says. This one is about who reads it, and the answer comes from my own nginx access logs.

The data

I have the site's HTTPS access logs from 2026-07-27 to 2026-09-17: 53 days and about 271,000 requests. I parsed every line, normalized the path, and pulled out every request for /llms.txt.

There were 162 of them. Here's who made them:

  • Chrome-Lighthouse: 147
  • curl: 9, all from my own IP address, on the day I created the file
  • Ordinary desktop browsers: 2, one of them also from my IP
  • Dataprovider.com's crawler: 2
  • A security scanner (OpenBash-Surface): 2
  • Any AI crawler or AI agent: 0

Who else was visiting

The zero would mean less if AI crawlers simply never came here. They do. In the same 53 days:

  • ClaudeBot: 5,698 requests
  • GPTBot: 4,060
  • Applebot: 2,603
  • Amazonbot: 2,164
  • meta-externalagent: 1,761
  • OAI-SearchBot: 427
  • Bytespider: 409
  • ChatGPT-User: 367

ClaudeBot alone fetched /robots.txt 734 times in that window. It checks the file that tells it what it may crawl, constantly. It never once asked for the file that was supposedly written for it.

Lighthouse, before and after

The Lighthouse numbers explain where the file came from. Lighthouse is what PageSpeed Insights runs, from Google's own servers, when someone audits a page. Of its 147 requests, 68 got a 404 and 79 got a 200. All 68 404s fall on a single afternoon: 2026-08-29, between 11:39 and 14:49. That was the afternoon I spent running PageSpeed Insights over and over, chasing its "Agentic Browsing" audit, and that audit failing on a missing llms.txt is exactly why I created the file.

At 14:54 my own curl got the first 200, and the commit adding the file (c6f24a1) landed at 14:54:21. Lighthouse kept requesting it for the rest of that afternoon and now got a 200 every time. After that, its requests thin out to a handful on scattered days in September. I can't tell from the logs who ran those audits.

So most of the file's 162 requests trace back to me: my own audits, my own curl checks, my own browser. In 53 days of logs, the audit that prompted the file is its only regular reader.

What this does and doesn't show

I want to be careful here, because "zero" invites more conclusions than the data supports.

It shows that none of the major AI crawlers fetched /llms.txt from this site over seven and a half weeks, while they were actively crawling the site's pages. That's a direct observation, not a survey.

It doesn't show that no AI system will ever read the file. A crawler could add support next month. A user could paste the URL into a chat. An agent could fetch it with a generic browser user agent that I'd count as "ordinary desktop browser". The two browser requests in the list could be anything. And these logs stop on 2026-09-17, so I can't speak to what's happened since.

It also doesn't measure whether the file helps anyone who does read it. It can't. All it can say is how many came.

The user-agent matching

My first pass at these numbers used a quick regex and came out slightly different: 159 requests and 6 from curl. This version parses each log line properly, matches the request path exactly rather than by substring (so requests for my articles about llms.txt don't count), and checks each user agent against named patterns for every AI crawler I know of, including ClaudeBot, Claude-User, GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Google-Extended, CCBot, Bytespider, Applebot, Amazonbot and meta-externalagent. The conclusion didn't move: zero, both ways.

Where that leaves the file

I'm keeping it. It's a few lines of Markdown, it costs nothing to serve, and it does pass the audit. But I've stopped thinking of it as a channel to AI systems. On the evidence of my own logs, it's an artifact of a Lighthouse check, read mostly by Lighthouse.

The file that AI crawlers verifiably read is the one they've always read: robots.txt. If I want to change how they treat this site, that's still where it happens.

Add new comment

Restricted HTML

  • Allowed HTML tags: <a href hreflang> <em> <strong> <cite> <blockquote cite> <code> <ul type> <ol start type> <li> <dl> <dt> <dd> <h2 id> <h3 id> <h4 id> <h5 id> <h6 id>
  • Lines and paragraphs break automatically.
  • Web page addresses and email addresses turn into links automatically.
Please share this article on your favorite website or platform.