llms.txt: we publish one, and no AI crawler has ever read it
We have published an llms.txt since May 2026. Across four days of IP-verified logs, AI crawlers fetched our robots.txt 120 times and our sitemap 43 times. They fetched llms.txt zero times.

There is a proposal called llms.txt: you put a markdown file at the root of your site, and it hands a language model a curated index of your best pages so the model does not have to crawl the whole site and guess what matters. A companion llms-full.txt holds the full text in one document.
We publish both, and have done since May 2026: prosopo.io/llms.txt is a hand-maintained 19 KB index, and llms-full.txt is 370 KB of full page text.
Nothing has ever read them.
What four days of logs say
prosopo.io sits behind a CDN that writes one access log line per request, with the client IP and user-agent. We took four days of those logs and counted requests to the three files a crawler might use to find its way around: robots.txt, the XML sitemaps, and llms.txt.
User-agent strings are no use here, since anyone can send any string they like, so we checked every request against the crawler IP ranges the vendors publish themselves — OpenAI's per-crawler JSON files, Perplexity's, DuckDuckGo's, and Anthropic's published AWS range — and counted only those from a verified address.
| File | Verified AI crawler fetches, 4 days |
|---|---|
/robots.txt | 120 — Anthropic 47, OpenAI 31, DuckAssist 27, Perplexity 15 |
/sitemap*.xml | 43 — Anthropic 40, OpenAI 3 |
/llms.txt and /llms-full.txt | 0 |
Zero. Not one verified AI crawler asked for the file built for it.
The files did get requests — 33 hits on llms.txt and 7 on llms-full.txt over the same period. Every one came from an address we could not match to any published crawler range. Some of that traffic asked for /blog/procaptcha-vs-hcaptcha-comparison/llms.txt, a path that has never existed. That is the mark of a scanner working through a wordlist, not an agent following a convention.
Why this is the expected result
This is not a scandal, and llms.txt is not a scam: it is an unratified community proposal from September 2024 which site owners have taken up, but the companies that run the crawlers have not.
OpenAI publishes the names and IP ranges of GPTBot, OAI-SearchBot and ChatGPT-User along with the robots.txt directives each one obeys, Anthropic does the same for ClaudeBot and Claude-User, and Perplexity, DuckDuckGo, Apple and Google all publish crawler names and robots directives. None of them documents llms.txt support, and our logs match that.
There is also a structural reason the file is a harder sell than it looks. A hand-curated index is the site owner's own claim about which of its pages are most valuable. That is the claim a search system exists to judge for itself, and it is easy to game. robots.txt works because it states a restriction — something only the site owner gets to decide. llms.txt states a recommendation, and nobody has to take it.
What is actually working
We have measured which of our pages AI assistants really fetch. The answer has nothing to do with llms.txt.
Over the same four days, verified AI crawlers fetched our pages 446 times, about 112 a day. The fetches cluster on one shape of page: analyst-report explainers, standards explainers, structured pricing comparisons. Our Forrester Wave breakdown, our Gartner and WAAP piece and our enterprise pricing comparison are fetched several times more often per page than anything else on the site.
Those pages all answer a comparative or factual question with structured, numeric, attributable data — named vendors, named analysts, named standards, real prices in a table. An assistant writing an answer needs facts it can quote and credit, and no index file replaces that.
What to do instead
If you have an afternoon to spend on being useful to AI systems, spend it in this order.
Get robots.txt right. It is the file they do fetch: 120 times in four days in our case. Decide which AI crawlers you want reading your content and name them. The distinction worth knowing is between training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended) and live retrieval agents (OAI-SearchBot, ChatGPT-User, Claude-User, PerplexityBot, DuckAssistBot). Block the first group and you stay out of training corpora. Block the second group and you drop out of answers today, which is almost certainly not what you want.
Keep your sitemap current. Fetched 43 times in four days, mostly by Anthropic. It is a solved problem and it still works.
Write the pages that get cited. Structured, factual, comparative, with the answer in the first paragraph and the numbers in a table rather than in prose.
Verify crawlers by IP, not user-agent. This matters more than the rest put together: get it wrong and your logs will tell you the wrong thing. In our four days, five Google Cloud addresses rotated fifteen separate AI crawler user-agents between them while requesting /.env, /.ssh/config, /actuator/threaddump and /api/env. Of their requests, 45 to 59% went to paths that have never existed on our site. If you had trusted the user-agent you would have concluded that Grok, Mistral and Cohere were crawling you hard, and every single xAI-branded request we saw came from that fleet rather than from a verifiable address.
And publish an llms.txt if you want. The convention may yet be adopted, and being early is cheap. Just do not put it in a report as a channel, and do not let it crowd out the robots.txt work. On today's evidence it is a file you maintain for an audience that has not arrived.
Our numbers, for reference
To reproduce this on your own domain you need CDN or server access logs with the client IP intact, and the published range files. Our one caveat is retention: our CDN keeps four days of logs and archives nothing, which is why this measurement covers four days rather than four months. If you want to track AI retrieval over time, turn on log archiving before you need it, not after.
Want to know which AI crawlers are hitting your site?
We can tell you which AI agents are fetching your pages, which are genuine and which are spoofing the user-agent. Tell us your domain.
Frequently Asked Questions
What is llms.txt?
llms.txt is a proposed convention, published at llmstxt.org in September 2024 by Jeremy Howard of Answer.AI, for a markdown file at the root of a website that gives large language models a curated index of its most useful pages. The idea is that an LLM working with a limited context window gets a hand-picked map instead of having to crawl and infer one. A companion file, llms-full.txt, carries the full text of those pages in one document.
Is llms.txt an official standard?
No. It is a community proposal, not an IETF or W3C specification, and no major AI provider has committed to reading it. OpenAI, Anthropic, Google and Perplexity all document the crawlers they operate and the robots.txt directives those crawlers obey. None of them documents llms.txt support.
Do AI crawlers actually read llms.txt?
In our own server logs, no. Over four days of access logs for prosopo.io we verified every request against the crawler IP ranges that OpenAI, Perplexity, DuckDuckGo and Anthropic publish. Those verified crawlers fetched /robots.txt 120 times and our XML sitemaps 43 times. They fetched /llms.txt and /llms-full.txt zero times. The only requests to llms.txt came from clients whose user-agent we could not verify.
Should I publish an llms.txt anyway?
It is cheap, so there is little reason not to, but treat it as a bet rather than a channel. It costs an afternoon to generate from your existing sitemap and it will do nothing measurable today. What demonstrably works right now is the boring pair: a robots.txt that names the AI crawlers you want to allow, and an XML sitemap that is actually current. Those are the two files the crawlers are provably fetching.
What does robots.txt have to do with AI crawlers?
It is the file they all actually read. GPTBot, OAI-SearchBot and ChatGPT-User, ClaudeBot and Claude-User, PerplexityBot, DuckAssistBot, Applebot-Extended and Google-Extended are each controllable by name in robots.txt, and each vendor publishes the IP ranges those crawlers originate from so you can verify them. If you want to shape how AI systems use your content, robots.txt is where the lever is.
How do I verify an AI crawler is genuine?
Check the IP, never the user-agent. OpenAI publishes JSON range files for each of its crawlers, Perplexity and DuckDuckGo do the same, and You.com is verifiable by reverse DNS. In our logs a fleet of five Google Cloud addresses rotated fifteen different AI user-agent strings while requesting /.env, /.ssh/config and /actuator — AI crawler names are a popular disguise for vulnerability scanning.
