ACADEMY
All posts

Does llms.txt Actually Get You Cited by AI? What Two 2026 Studies Found

Two 2026 studies tracked llms.txt across more than 400,000 domains combined. Neither found a link to AI citations, and the server logs show why.

If you've spent any time in SEO or GEO circles in the last year, you've probably been told to add an llms.txt file to your site "so AI can find your content." It sounds like reasonable advice, robots.txt tells crawlers what not to touch, sitemap.xml tells search engines what exists, so a file that tells AI models what your site is about seems like the obvious next step. Two large studies published in 2026, one from Ahrefs covering 137,210 domains and one from SE Ranking covering roughly 300,000, both looked at whether llms.txt actually does anything for AI visibility. The data says no, and the reason why is more useful than the file itself.

What llms.txt actually is

llms.txt isn't a protocol any browser or crawler is required to honor, it's a convention. Jeremy Howard, co-founder of Answer.AI and fast.ai, proposed it in September 2024 as a plain Markdown file served at a domain's root. The format is deliberately minimal: one required H1 with the site name, an optional one-paragraph summary in a blockquote, and then H2-delimited lists of links to the pages a model should read first, each with a short description.

The original problem it was trying to solve was narrow and real: a model reading raw HTML wastes tokens on navigation, ads, and markup, and a full site is bigger than any context window. llms.txt hands an AI system a curated, pre-cleaned index instead of making it parse a page itself, similar in spirit to what we cover in why AI agents get slower and less accurate the longer their context grows. That's a legitimate mechanism. What it was never designed to do, despite how it's been marketed since, is influence whether an AI search engine cites your site in an answer.

The adoption number nobody expected

SE Ranking's analysis of nearly 300,000 domains, published in November 2025, found that only 10.13% had adopted an llms.txt file at all, roughly nine in ten sites hadn't bothered. More importantly, adoption didn't cluster where you'd expect. High-traffic sites (100,001+ monthly visits) had the lowest adoption rate of any traffic tier, at 8.27%, while mid-traffic sites adopted it more often, at 10.54%. If llms.txt were actually driving AI visibility, you'd expect the sites competing hardest for AI citations to adopt it first. They didn't.

The correlation test is the part that matters most. SE Ranking ran statistical correlation tests plus an XGBoost regression model against AI citation frequency across major LLM answers. Removing the llms.txt variable from the model improved its predictive accuracy, meaning the feature wasn't just neutral, it was adding noise. Their own conclusion: llms.txt "doesn't seem to directly impact AI citation frequency. At least not yet."

Who's actually requesting the file

Ahrefs' June 2026 study went further and looked at server logs instead of just presence or absence. Across 137,210 domains, 28% had published an llms.txt file, and of those, 97% received zero recorded requests in the entire month of May 2026. Of the roughly 38,000 files that were technically live, only about 1,100 got any traffic at all.

The breakdown of who that traffic actually came from is the most useful part of the whole study, because it kills the assumption that the requests are even coming from AI systems in the first place.

Who actually requests your llms.txt file — share of requests by source category. Source: Ahrefs, 137,210 domains analyzed, June 2026
Who actually requests your llms.txt file — share of requests by source category. Source: Ahrefs, 137,210 domains analyzed, June 2026

SEO audit tools alone generated more requests (21.7%) than every category of AI bot combined (19.5%). AI assistants like ChatGPT and Claude, the systems people are actually trying to get cited by, accounted for just 2.5% of requests. Most of the traffic hitting your llms.txt file is other SEO tools checking whether you have one, not models reading it to decide what to cite.

What the two studies actually measured

StudyScopeMethodHeadline finding
SE Ranking (Nov 2025)~300,000 domainsCorrelation + XGBoost regression against AI citation frequencyNo measurable link between llms.txt presence and citations; removing the variable improved model accuracy
Ahrefs (Jun 2026)137,210 domains, server-log levelDirect request-log analysis of who fetches llms.txt files97% of files get zero requests; AI bots make up under 20% of the traffic that does arrive

Two independent methodologies, one on citation outcomes and one on actual server traffic, landed on the same conclusion from different directions. That's a stronger signal than either study alone.

What the platforms themselves say

This isn't just third-party research contradicting a vendor claim, the platforms have said it directly. At Google's Search Central Deep Dive in July 2025, Gary Illyes stated that Google does not support llms.txt and has no plans to, and that ranking in AI Overviews depends on the same fundamentals as regular SEO. Google's John Mueller has gone further, describing llms.txt as "not done for search" and calling it, at best, "a temporary crutch, perhaps to save some tokens" for coding assistants specifically, not a discovery mechanism for AI search.

That framing matches where llms.txt actually earns its keep: developer documentation. Anthropic serves an llms.txt for its developer platform docs, and Vercel and Stripe do the same for theirs. In that context the audience isn't a search engine deciding whether to cite you, it's a coding assistant like Claude Code or Cursor trying to write correct integration code without hallucinating an endpoint that doesn't exist. That's a real, narrow win. It's just not the win most sites installing llms.txt think they're getting.

What actually moves AI citations

If llms.txt isn't the lever, what is? The same evidence that discredits llms.txt keeps pointing back to fundamentals. Our breakdown of Ahrefs' 1,885-page schema markup study found the same pattern: structured data alone barely moved citation rates either. What actually correlates with getting cited is ranking well organically in the first place, since AI Overviews and chat answers still lean heavily on the same index Google Search already ranks. If you haven't read how the underlying goal shifted from ranking to being quoted, GEO vs SEO: what actually changes when AI answers the question is the right starting point before you spend engineering time on any AI-specific file format.

The practical takeaway

Skip llms.txt unless you run a developer platform where coding assistants are a meaningful chunk of your audience, in which case it costs almost nothing to ship. For everyone else chasing AI citations, the two 2026 studies point the same direction: put the effort into content that already ranks, that answers the question directly in the first two sentences, and that a model can verify against a source it already trusts. That's slower than adding a text file, but it's the version that's actually shown up in the data. If you want the full framework for building content AI search engines cite instead of chasing format tricks, that's exactly what we walk through inside The AI Marketing Engine.

Go deeper

The AI Marketing Engine

Courses launching soon

Want this as a full course, not just a post?

Join the waitlist and get a 20% launch discount the moment we open checkout. No payment now.

Or join our free community for AI tips while you wait