# What an AI Agent Sees on 1,567 WordPress Sites in 2026: Thin Text Costs More Than robots.txt

> We probed 1,567 WordPress sites as an agent on 25 September 2026. robots.txt rules blocked 48 of them. Having nothing readable on the page lost 157.

- Published: 2026-09-25
- Updated: 2026-10-04
- Author: xSpeed Cache Team
- Tags: AI, AI Search, Research, Performance, WordPress, MCP, check:D1, platform:wordpress
- Canonical: https://xspeedcache.com/blog/what-ai-agents-see-wordpress-site/

---

Updated September 2026

We fetched the homepage of 1,567 WordPress sites on 25 September 2026, twice each: once as a browser, once as something plainly not a browser. Then we counted the words an agent could read without running any JavaScript. The median site offered 746 of them. One in ten offered fewer than 200.

The comparison people expect here is robots.txt against the AI crawlers. The more useful comparison is permission against substance, because those two gates cost very different numbers of sites. Blocking rules cost 48 sites in our sample. Having nothing readable on the page cost 157.

This is a study of what reaches an agent, measured on one page per host. It is not a study of whether an agent then chose to cite anybody.

## Quick Summary: Which of the Four Gates You Are Losing

| If you want to know… | Do this | Why it works |
|:---|:---|:---|
| whether you are blocked at all | `curl -s https://yoursite.com/robots.txt \| grep -iE 'gptbot\|claudebot\|ccbot'` | 94.6% of the sites we measured name no AI agent at all, so most owners have never made this decision |
| whether your origin serves non-browsers | fetch your own homepage with a non-browser user agent and compare the status code | 39 sites served a browser and refused us |
| whether bytes arrive in time | read the first-byte figure on any speed report | 29.1% of sites took longer than a second to start answering |
| whether there is anything to read | strip the tags from your own HTML and count the words | this gate lost more sites than the other three together |
| whether an agent has a shortcut | request `/llms.txt` and check the content type | 30.3% already serve one, and 66 sites serve an HTML page pretending to be one |

## The Four Gates, and Why the Order Matters

An agent reading your site passes four tests in sequence. Each one is cheap to check, each one is binary, and a site that fails an early gate never reaches the later ones. That ordering is the whole point: optimising gate four while failing gate one buys nothing.

| Gate | Question | How we measured it | Sites lost |
|:---|:---|:---|---:|
| 1. Permission | does robots.txt allow a named AI agent? | parse every group, test whether `Disallow: /` applies | 48 |
| 2. Admission | does the origin answer a non-browser request? | one fetch as a browser, one as our own identified probe | 127 |
| 3. Budget | do the bytes start arriving quickly? | time to first byte on our own fetch, 3s threshold | 41 |
| 4. Substance | is there readable text in what arrived? | strip tags, count word-like tokens, 200-word threshold | 150 |

Run in that order, 1,201 of 1,567 sites clear all four. That leaves 366 sites, 23.4% of the sample, losing an agent somewhere. Of those, 354 fail exactly one gate and 12 fail two. None fails three.

![Funnel chart: four gates applied in sequence across 1,567 WordPress sites, leaving 1,201](https://xspeedcache.com/images/blog/agent-gates-funnel.webp)

## What This Study Cannot Tell You

The method comes before the results, and so do its limits. Five of them matter enough to state plainly.

**We did not impersonate anyone.** Our probe identifies itself as `xSpeedScan-AgentProbe/1.0`, not as GPTBot or ClaudeBot. So gate 2 measures whether an origin will serve a user agent that is not a browser. It does not measure whether that origin blocks a specific named crawler, and a site that allowlists three bots by name registers here as refusing us while serving them perfectly well.

**We did not execute JavaScript.** Our own harness could not drive headless Chromium against remote HTTPS in this environment, so every word count below is the text present in the HTML as delivered. That is exactly what a non-rendering reader gets, and it is not what a rendering reader gets. We cannot tell you how much text JavaScript would have added.

**One URL per host, and it is the homepage.** A homepage is usually the best-maintained page on a site and often the thinnest in prose. Both biases are real and they push in opposite directions.

**The 3-second budget is ours, not anybody's published figure.** No crawler operator we could find publishes a timeout, so we report the whole curve rather than only the verdict.

**The sample is self-selected.** These are sites somebody chose to run through our own scanner, which skews toward sites with a suspected speed problem, and 745 of the 1,567 already run our plugin. This is not a random sample of WordPress.

## Gate 1: The Gate Everyone Guards and Almost Nobody Uses

robots.txt returned 200 on 1,415 of the 1,567 hosts, or 90.3%. The remaining 152 either timed out, 404ed or refused us. Then the interesting part: only 84 sites, 5.4%, name any AI agent at all. Forty-eight block at least one.

| Agent | Sites naming it | Sites blocking it |
|:---|---:|---:|
| GPTBot | 68 | 30 |
| Google-Extended | 67 | 33 |
| ClaudeBot | 57 | 24 |
| CCBot | 51 | 39 |
| PerplexityBot | 51 | 13 |
| Applebot-Extended | 44 | 31 |
| OAI-SearchBot | 39 | 13 |

Two things fall out of that table. CCBot is the most blocked and only the fourth most named, which is what a rule copied from a blog post in 2023 looks like years later. And the search-facing agents are blocked far less often than the training-facing ones: OAI-SearchBot draws 13 blocks against GPTBot's 30, which suggests owners who understood that OpenAI treats the two independently. OpenAI's own [crawler documentation](https://platform.openai.com/docs/bots) confirms the split, describing a webmaster who allows OAI-SearchBot for search results while disallowing GPTBot for model training.

Anthropic [documents three robots](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), ClaudeBot, Claude-User and Claude-SearchBot, and states that its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt". Per the same page they also respect anti-circumvention measures and will not attempt to bypass a CAPTCHA. That last sentence is why gate 2 exists.

## Gate 2: Whether Your Origin Answers Anything That Is Not a Browser

Our agent user agent got a 200 from 1,442 hosts, 92.0%. The browser user agent got one from 1,426. More sites served the agent than the browser, which is the first sign that this gate is not mostly about bot blocking.

Breaking the 125 non-200s apart settles it:

- **86 hosts refused both** user agents, or did not answer at all. These sites are down, misconfigured or gone. An agent's experience of them is no worse than a person's.
- **39 hosts served the browser and refused us.** This is the real UA-discrimination number, 2.5% of the sample. Twelve returned 403; the other 27 dropped the connection.
- **15 more served a 200 that was a challenge page**, not content. "Just a moment" is a 200, and a naive checker counts it as a success.

That last trap is worth dwelling on, because it is invisible from a status code alone. An agent that respects anti-circumvention policy, as Anthropic's documentation says its bots do, reads that page, finds no content, and leaves. The site owner sees a 200 in the log.

## Gate 3: The Budget Nobody Publishes

Time to first byte, measured on our own fetch of the 1,442 hosts that served us, from a single probe location. The median host started answering in 619ms, the 75th percentile in 1,101ms and the 90th in 1,764ms. The tail is what decides this gate:

| Slower than | Sites | Share |
|:---|---:|---:|
| 1 second | 420 | 29.1% |
| 2 seconds | 103 | 7.1% |
| 3 seconds | 41 | 2.8% |
| 5 seconds | 12 | 0.8% |

Read the curve rather than our threshold. At three seconds this gate costs 41 sites and looks minor. At one second it touches 420, nearly three in ten. Where the real line sits depends on a timeout nobody publishes, which is why we report all of it. Our archive's own probe, measuring from Vilnius with a warm cache, puts the median at 310ms, about half our figure, and the difference between those two numbers is a fair estimate of how much measurement location alone moves this gate.

This is the one gate a caching plugin directly moves, and it is also the gate that costs the fewest sites at any reasonable threshold. We publish that in the same breath because both halves are true.

![Bar chart of time to first byte across 1,442 WordPress hosts, median 619ms](https://xspeedcache.com/images/blog/agent-gates-ttfb.webp)

## Gate 4: The Gate That Actually Bites

Of the 1,440 hosts that served a real page, here is the distribution of readable words in the delivered HTML:

| Percentile | Words |
|:---|---:|
| 10th | 140 |
| 25th | 388 |
| median | 746 |
| 75th | 1,282 |
| 90th | 2,113 |

And the thresholds: 79 sites under 50 words, 111 under 100, 191 under 200, 279 under 300, 474 under 500. A third of WordPress homepages in this sample deliver under 500 readable words to a reader that does not run JavaScript.

Two expected findings did not appear, and both are worth stating because they contradict the usual advice. **Not one of the 1,440 sites showed a JavaScript shell marker** in its HTML: no empty React root, no Nuxt or Next hydration payload standing in for content. WordPress renders on the server, so the single-page-application failure that dominates this conversation elsewhere is essentially absent here. And exactly **one site of 1,440 served materially different text** to our agent user agent than to our browser one. Cloaking, as a measurable phenomenon on WordPress, does not exist at this sample size.

The [llms.txt proposal](https://llmstxt.org/) names the underlying problem precisely. Per its v2 text, authored by Jeremy Howard and last modified 10 August 2026: "An HTML page wraps its information in navigation, ads, and JavaScript, and converting it back into clean text is difficult and imprecise."

## The Thin-Text Audit: What Those 191 Sites Actually Were

A cohort defined by a threshold is a cohort of whatever happens to sit under it, so we refetched all 191 and classified them. This is the step that turns a number into a finding.

| What it turned out to be | Sites | Counted as a gate-4 failure? |
|:---|---:|:---:|
| Genuinely thin real site | 157 | ✅ |
| Default WordPress install, "Hello world!" | 15 | ❌ |
| Placeholder or coming-soon page | 12 | ❌ |
| No longer serving 200 on refetch | 6 | ❌ |
| Challenge page our first pass missed | 1 | ❌ |

So the defensible gate-4 number is **157 of 1,440 served pages, 10.9%**, not the 191 the raw threshold produced. The audit removed 34 sites, and one of them exposed a bug in our own challenge detector, which we are reporting rather than quietly fixing.

The 157 are ordinary sites. A print shop whose homepage is 279KB of HTML carrying 89 words, because the products are images and the copy lives in the image. A consultancy whose homepage is a nav bar and a tagline. None of them is broken. All of them are invisible to a reader that cannot see pictures, and that is a content decision rather than a bug.

Google is the exception that proves the shape of this. Per its [JavaScript SEO documentation](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics), "Google processes JavaScript web apps in three main phases: Crawling, Rendering, Indexing", and Googlebot queues pages for rendering as a distinct step. Neither OpenAI's crawler overview nor Anthropic's crawler page mentions JavaScript anywhere. We checked both documents on 25 September 2026: zero occurrences. That asymmetry is documented rather than inferred, and until it changes, planning for text that is present in the HTML is the only safe assumption.

## llms.txt Arrived on WordPress by Plugin, Not by Decision

475 of the 1,567 hosts, 30.3%, serve a real `/llms.txt`: 200 with a plain-text content type and the format's `#` heading. That number is far higher than the 5.4% who have touched robots.txt for an AI agent, and the reason is visible in the files themselves.

| Generated by | Files |
|:---|---:|
| Yoast SEO | 119 |
| Rank Math | 90 |
| All in One SEO | 48 |
| Unattributed | 217 |
| Other plugin | 1 |

258 of the 475, 54.3%, carry a generator line naming an SEO plugin. Those site owners did not decide to publish an llms.txt. They updated a plugin. Median file size is 4,892 bytes.

A separate 66 sites, 4.2%, return a 200 with an HTML content type at that path. That is a soft 404: an agent asking for a clean text index gets a rendered page claiming success. Check the content type, not the status code.

## Running the Four Gates Across 25 Client Sites

Checking one site takes four commands. Checking twenty-five by hand is an afternoon, which is why nobody does it twice. The free xSpeed Scan grades the delivery and speed half of this on any public URL without an account, and every report lands in the [public scan archive](https://xspeedcache.com/scan/all/) that supplied this study's corpus. Its check D1 is gate 3 under another name.

For the agent-readiness half, gates 1, 2 and 4, hand off to [AIScan](https://aiscan.site/), which is ours too and grades exactly that: robots.txt rules for named AI agents, whether an agent gets served, llms.txt, and whether the content survives the trip. Covering half of agent readiness in a caching article and calling it done would be the worse choice. For the content side of the same question, our guide to [visibility in AI search](https://xspeedcache.com/blog/boost-your-wordpress-websites-visibility-in-ai-search/) covers what a cited page looks like once an agent can read it.

If you would rather automate it, our [MCP server documentation](https://xspeedcache.com/docs/scan-mcp/) covers driving the scan from Claude or any MCP client, and the [AI agents use case](https://xspeedcache.com/use-cases/ai-agents/) walks the fleet version. Ready-made prompts are in [Scan site speed with Claude](https://xspeedcache.com/agent/claude/speed-scan/), and [AI agents](https://xspeedcache.com/agent/) covers the other clients. For the wider category, our [roundup of WordPress MCP servers](https://xspeedcache.com/blog/best-wordpress-mcp-servers-2026/) re-queries the directory monthly, and the [two performance servers compared head to head](https://xspeedcache.com/blog/xspeed-cache-mcp-vs-wp-rocket-mcp/) sets ours against WP Rocket's tool by tool. Hosting sets the floor under gate 3 that no plugin can lift: we recommend [xCloud](https://xcloud.host/) for it, and it is ours, since xCloud and WPDeveloper are both Startise companies. Which server you land on also decides what a cache can do at all, which our [server compatibility notes](https://xspeedcache.com/docs/server-compatibility/) set out.

## What Our Own Users' Sites Score, Including Where We Lose

745 of these 1,567 sites run xSpeed Cache and 822 do not. Splitting the sample that way is the fairest test available to us, and we publish both columns.

| Measure | xSpeed sites (745) | Everyone else (822) |
|:---|---:|---:|
| Median time to first byte | 628ms | 613ms |
| Slower than 3s | 2.5% | 3.1% |
| Median readable words | 687 | 776 |
| Genuinely thin (under 200 words) | 12.1% | 9.8% |
| Serves a real llms.txt | 29.7% | 30.9% |
| Blocks a named AI agent | 3.2% | 2.9% |

We come out marginally ahead on the slow tail and marginally behind on median first byte. Both gaps are small enough that this study cannot attribute them, and our own benchmark work on [why a single scan's score moves](https://xspeedcache.com/blog/wordpress-caching-plugin-benchmark-protocol/) is the reason we will not try.

On gate 4 we are clearly behind: 12.1% of our users' sites are genuinely thin against 9.8% of everyone else's, and their median page carries 89 fewer readable words. A page cache does not write your copy, and sites that install a cache plugin are disproportionately sites that already had a speed problem, a population we measured directly in our [cache headroom study](https://xspeedcache.com/blog/when-not-to-use-caching-plugin/). Both facts push that column the way it went. We would rather print it than explain it away, and it marks the honest edge of what caching is for: gate 3 is ours to fix, and gates 1 and 4 are yours.

That boundary is the same one our [cache ceiling study](https://xspeedcache.com/blog/caching-did-not-fix-slow-wordpress/) found from the performance side, where subtracting the entire server response left 93.4% of failing sites still failing. Our plugin is at version 1.3.5 with 7,000 active installs and 5.0 from 14 ratings, per a [wordpress.org Plugin API](https://wordpress.org/plugins/xspeed/) query on 25 September 2026. Against LiteSpeed Cache's seven million, we are the newcomer here, and the 745-site sample above is our own user base rather than the field.

## Common Mistakes Site Owners Make at These Gates

- **Auditing gate 1 and stopping.** It is the cheapest gate to check and the one that cost the fewest sites, 48 against gate 4's 157.
- **Trusting a 200.** Fifteen sites in this sample returned 200 with a challenge page and 66 returned 200 with an HTML file at `/llms.txt`. Read the content type and the first line of the body.
- **Copying an AI blocklist without rereading it.** CCBot leads the block list and trails the naming list, which is what a stale copied rule looks like.
- **Adding an llms.txt while the HTML stays thin.** The file points at pages. If the pages carry 89 words, a tidier index does not help.
- **Treating this as an SEO task.** Gates 2 and 3 are delivery, and gate 4 is often image-heavy design. Our guide to [reading a speed report](https://xspeedcache.com/blog/read-pagespeed-report-wordpress/) covers which number to act on first.
- **Blocking a search-facing agent by accident.** GPTBot and OAI-SearchBot are independent settings per OpenAI's documentation, and 13 sites here block the one that produces referral traffic.

## Frequently Asked Questions

### I blocked GPTBot last year. Did I also block ChatGPT from citing my site?

Not necessarily. Per OpenAI's crawler documentation the tags are independent: GPTBot governs training, OAI-SearchBot governs appearing in search results, and ChatGPT-User covers a page fetched because a user asked. Thirteen sites in our sample block OAI-SearchBot, which is the one tied to referral traffic. Note also that OpenAI documents a lag of roughly 24 hours before a robots.txt change takes effect.

### My homepage looks full of text in a browser. Why did you count 89 words?

Because the words were in the pictures. That was a real site in our 157, a print shop serving 279KB of HTML for 89 readable words. Strip the tags from your own HTML and count what is left. If the gap is large, the copy lives in images.

### I have no robots.txt at all. Is that a problem for agents?

It is the default, not a problem. 152 of our 1,567 hosts served no readable robots.txt and 94.6% of the rest name no AI agent, so an absent file puts you with the majority and blocks nothing. It does mean you have made no decision either way.

### Does my site need to run JavaScript-free to be readable?

On WordPress you are almost certainly fine already. Not one of the 1,440 pages we measured showed a JavaScript shell marker. The thin-text problem we found is images and short copy, not client-side rendering.

### I installed a caching plugin and my agent visibility did not change. Why?

Because a cache moves gate 3 and nothing else. Our own users' sites sit at 2.5% over the three-second line against 3.1% for everyone else, a small win, while their median readable word count is 89 words lower. Caching is a delivery fix, and three of these four gates are not delivery.

### Should I add an llms.txt?

It is cheap and 30.3% of the sites we measured already serve one, usually because Yoast, Rank Math or All in One SEO generated it on an update. Check yours returns a plain-text content type rather than HTML, because 66 sites here fail that test while returning 200.

### Is a 403 to a crawler ever the right answer?

Sometimes, and it is rarer than people think: 39 of 1,567 sites served a browser and refused our non-browser probe. If that is deliberate, you know what you are trading. If it came with a security plugin's default, you are refusing readers you wanted.

### My host is fast for me and slow in your table. Which is right?

Both, and the gap is the point. Our archive's probe in Vilnius reports a 310ms median where our probe reports 619ms on the same population. Measurement location moves this gate by roughly a factor of two, which is why we published the curve.

### How do I check all four gates without running a study?

Four commands per site: fetch robots.txt, fetch the homepage with a non-browser user agent, read the first-byte figure, and count the words after stripping tags. For more than a handful of sites, our [free vs Pro comparison](https://xspeedcache.com/free-vs-pro/) shows where the fleet tooling starts.

### Does a CAPTCHA or bot-check page block AI agents outright?

For the ones that document a policy, yes. Anthropic's crawler page states its bots will not attempt to bypass CAPTCHAs on sites they crawl. So a challenge page is a stop rather than a slow path, and it still logs as a 200 on your side.

### Is 200 readable words the right threshold for gate 4?

It is ours, chosen because it sits near the 10th percentile of this sample. The distribution matters more than the line: 79 sites are under 50 words and 474 are under 500, so pick the threshold that matches what your pages are supposed to say.

![Distribution of readable words without JavaScript, and the thin-cohort audit](https://xspeedcache.com/images/blog/agent-gates-words.webp)

## Conclusion: Fix the Gate You Are Actually Losing

Three in four sites in this sample clear all four gates. Among the 366 that do not, 354 fail exactly one, which means almost every site with an agent problem has a single, identifiable, fixable cause rather than a general condition.

| If your goal is… | Check first | Expected finding |
|:---|:---|:---|
| stop being excluded outright | robots.txt for named agents | you probably name none, and block none |
| be readable by a non-browser | status code and body under a non-browser UA | 97.5% of sites are already fine |
| arrive inside a timeout | first-byte time, against the whole curve | 29.1% are over a second |
| have something worth reading | word count after stripping tags | the gate that lost the most sites |

**What to do this week:** run one non-browser fetch of your own homepage and read the status code. Strip the tags and count the words, then compare that number to what you think the page says. Check whether `/llms.txt` returns plain text or an HTML page wearing a 200. Scan the delivery half at [xSpeed Scan](https://xspeedcache.com/scan/) and the agent-readiness half at AIScan, both free and neither needing an account. If the four gates turn into an ongoing job across a fleet, xSpeed Cache is free on wordpress.org and Pro starts at $29 a year at founding pricing with a 14-day money-back guarantee, and the [80-capability comparison](https://xspeedcache.com/comparison/) shows exactly where it sits against the field.

If you run this on your own sites, the four numbers are comparable to ours and we would like to see them. The measurement is four commands, the corpus is public, and the part worth arguing about is where the thresholds belong.
