# The Smallest Score Gap Was the Worst News: Triaging 1,110 WordPress Speed Reports Graded on Two Devices in 2026

> 1,110 WordPress pages graded on desktop and mobile on 30 September 2026. The letter grade differed on 61.6%, and 8 of 19 checks never moved at all.

- Published: 2026-09-30
- Updated: 2026-10-04
- Author: xSpeed Cache Team
- Tags: Core Web Vitals, Mobile, PageSpeed, Performance, Research, platform:wordpress
- Canonical: https://xspeedcache.com/blog/mobile-desktop-score-divergence/

---

Updated September 2026

Our own scan archive graded 1,110 WordPress pages twice on 30 September 2026, once as a desktop run and once as a mobile run of the same URL, and the letter grade came out different on 61.6% of them. The median difference was 12.0 points. That much is a measurement quirk, and it is already documented: per [Google's PageSpeed Insights documentation](https://developers.google.com/speed/docs/insights/v5/about), the two runs use different hardware and different networks, so they were never going to agree.

The comparison people expect from two device columns is which one is right. The more useful comparison is what the difference between them rules out, because 8 of the 19 checks in the report never moved between the two runs at all. This piece measures which parts of a speed score follow the device and which parts do not, on one archive, on one day.

## Quick summary: which of the two numbers to act on

| If your report shows… | Do this | Why |
|:---|:---|:---|
| A wide gap, mobile far below desktop | Work on JavaScript execution and image bytes | The server side scored identically in both runs, so it is not the cause |
| A narrow gap and a low score in both | Work on the server first: cache, compression, first byte | 18.8% of sites sit here, with 161 ms more first-byte latency than the wide-gap group |
| A narrow gap and a high score in both | Nothing | 10.9% of sites, median delivery score 78 |
| Any gap under about 12 points | Treat it as noise, not a signal | A single lab run moves that much on its own, [as we measured on one unchanged page](https://xspeedcache.com/blog/wordpress-caching-plugin-benchmark-protocol/) |

## The two views are one scan, and only one of them is graded

A scan report on [xSpeed Scan](https://xspeedcache.com/scan/) carries two graded views of a single visit. The server was probed once. The page was then rendered twice by Google's Lighthouse, once under each form factor, and each view scores the same twenty-check rubric against its own run.

Google states the difference plainly. Per [the PageSpeed Insights documentation](https://developers.google.com/speed/docs/insights/v5/about), *"Lighthouse simulates the page load conditions of a mid-tier device (Moto G4) device on a mobile network for mobile, and an emulated-desktop with a wired connection for desktop."* One page, two machines, two networks.

Which view a reader sees first matters, and we checked rather than assumed it. Recomputing each report's headline score from its own four dimension totals reproduced the published number on **1,890 of 1,890 reports**, and on every report that carries a desktop view, **1,129 of 1,129**, that published headline is the desktop run. Google ranks the other one: per [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/mobile/mobile-sites-mobile-first-indexing), *"Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking."*

## The method, and the seven things it cannot tell you

Every scan report URL was fetched from `https://xspeedcache.com/scan/r/sitemap.xml` on 30 September 2026: 2,870 reports, of which 2,869 parsed and one connection dropped. 1,895 are WordPress. A report entered the study only if it carried all four dimension scores in **both** views, which leaves **1,110 WordPress hosts**, every one of them a distinct hostname.

The limitations come before the results, because two of them bound what the rest is worth.

1. **The gap inherits the instrument's noise, and the noise is about the size of the median gap.** Scanning one unchanged page four times in twelve minutes returned 70, 61, 62 and 73. The median gap here is 12.0 points. So the split below is a direction, not a measurement, and no single site's gap should be read to the point.
2. **One check inside the server dimension is not a server check.** Check D4 scores the cache policy on static assets from a Lighthouse audit rather than a probe, which is why the delivery dimension is not perfectly frozen. It is named again in the results.
3. This is observational. Sites that get scanned chose to be scanned, so the archive is not a sample of WordPress.
4. One URL per host, almost always the homepage.
5. Lab, not field. [Google's own documentation](https://developers.google.com/speed/docs/insights/v5/about) says lab data *"may not capture real-world bottlenecks."*
6. Server probes run from Vilnius, so every latency figure carries network distance and reads high.
7. The two-view report is recent. 1,415 reports carry both views; the rest of the archive predates it.

## Result 1: Eight of nineteen checks never moved

Nineteen of the twenty rubric checks scored points on enough sites to compare. Eight of them returned **exactly the same points in both views on every single site**, with no exceptions: first byte, page served from cache, HTML compression, redirect chain, CDN presence, and the three platform checks.

![Movement rate for each rubric check between the desktop-graded and mobile-graded run of the same URL, from 95.9% on the Lighthouse performance score down to eight checks at zero.](https://xspeedcache.com/images/blog/two-view-split-check-movement.webp)

Two rows in that chart are worth stopping on.

**D4 is the reason the delivery dimension is not clean.** It moved on 20.9% of sites, and it is the only check in that dimension that moved at all. It reads a Lighthouse audit of static-asset cache headers, so it inherits the device. A dimension labelled "Server response and caching" is therefore device-dependent on a fifth of sites, which is a scanner design wrinkle rather than a fact about anyone's server.

**Two checks moved against the phone's favour.** Cumulative Layout Shift scored worse on desktop on 10.2% of sites against 7.7% on mobile, and total page weight scored worse on desktop on 9.9% against 3.2%. The wider desktop viewport requests larger candidates from a `srcset`, so the desktop run genuinely downloads more bytes. Image format scoring moved on 32.0% of sites and went both ways for the same reason.

## Result 2: The whole gap lives in one dimension

Splitting each site's gap across the four dimensions, on the normalised hundred-point scale:

| Dimension | Mean contribution to the gap | Share of the mean gap | Moved at all | Identical in both views |
|:---|:---:|:---:|:---:|:---:|
| Core Web Vitals and lab metrics | +11.37 pts | **96.6%** | 96.8% | 4.05% |
| Server response and caching | +0.31 pts | 2.7% | 20.9% | 79.10% |
| Asset optimization | +0.29 pts | 2.5% | 54.6% | 39.82% |
| Platform readiness | +0.00 pts | 0.0% | 0.0% | **100.00%** |

The mean total gap is 11.77 points and the median is 12.0. Desktop scored higher on 92.3% of sites, the two views tied on 2.1%, and **mobile scored higher on 5.7%**. On the raw Lighthouse performance score the median gap is 16 points and mobile wins on 9.0%, so the minority where the phone does better is real and not a rounding artifact.

## The Two-View Split: reading your own report in ninety seconds

The framework is a subtraction, and it needs no tooling beyond a report you already have.

1. **Read the mobile score.** That is the floor. It is the number Google's indexing uses, and no device swap improves it.
2. **Subtract it from the desktop score.** The difference is the portion of your score that exists only because the test hardware changed.
3. **Read the delivery and platform rows in either view.** They are the same in both by construction, on 100.00% and 79.10% of sites respectively. Whatever they cost you is charged on every device.
4. **Decide from the pair, never from the gap alone.** A wide gap with a decent mobile score is a good position. A narrow gap with a poor score in both is the worst one.

What the split buys you is an exclusion. If your delivery row is identical in both views, then the gap cannot have come from your host, your cache or your compression, and no hosting change will close it. That is worth knowing before anyone buys anything.

Its limits, stated plainly: the gap is noisy below roughly 12 points, it says nothing about which script or which image is responsible, and it is a lab result that a field measurement may contradict.

## Result 3: The grade changes on six sites in ten

Letter grades differed between the two views on **684 of 1,110 sites, 61.6%**. The direction is almost entirely one way: 661 sites grade lower on mobile and 23 grade higher. The largest cells are 256 sites reading D on desktop and F on mobile, 228 reading C then D, 78 reading C then F, 60 reading B then C, and 26 reading B then D.

The practical version: a reader who opens the report, sees a C, and closes it has a better than even chance of having read a grade the ranking device does not give.

## Result 4: The smallest gaps hide the slowest servers

Sorting every site by whether its Largest Contentful Paint passed in each view produces four states, and three of them are populated.

![Three cohorts by LCP verdict. Sites failing on both devices have a median TTFB of 428 ms and a gap of 8.1 points, against 267 ms and 15.3 points for sites failing on mobile only.](https://xspeedcache.com/images/blog/two-view-lcp-triage-cohorts.webp)

The cohort failing on both devices has the **slowest** server in the study, a median first byte of 428 ms against 267 ms for the mobile-only failures, and the **smallest** gap, 8.1 points against 15.3. Its delivery score is 62 against 72, and its asset score is 20 against 38. One site in 1,110 failed on desktop alone.

The same inversion holds when sites are grouped by gap size instead of verdict.

![As the desktop-to-mobile gap widens across four bands, median mobile LCP rises from 2.5s to 9.0s while median TTFB falls from 359ms to 208ms.](https://xspeedcache.com/images/blog/score-gap-ttfb-inversion.webp)

| Gap band | Sites | Median TTFB | Median delivery score | Median mobile LCP | Mobile LCP failing |
|:---|:---:|:---:|:---:|:---:|:---:|
| Under 5 pts | 223 | 359 ms | 70 | 2.5 s | 50.2% |
| 5 to 10 pts | 225 | 367 ms | 68 | 4.8 s | 96.0% |
| 10 to 20 pts | 495 | 267 ms | 72 | 7.7 s | 99.6% |
| 20 pts or more | 167 | **208 ms** | **75** | **9.0 s** | 100.0% |

Spearman's rank correlation between the gap and mobile LCP is **+0.405** across 1,105 hosts. Between the gap and first-byte time it is **−0.069**, which is nothing. A wide gap is a signal about the browser, and a mild signal that the server work is already finished.

This does not license a causal claim. These are sites that differ in many ways at once, and the association is an ordering rule for where to look first, not a promise about what a fix will return. It is the same caution that applies to [the cache ceiling](https://xspeedcache.com/blog/caching-did-not-fix-slow-wordpress/) and to [the 350 KB page-weight threshold](https://xspeedcache.com/blog/page-weight-vs-lcp-wordpress/): archive data orders the work, it does not guarantee the outcome.

## Our own cohort, in both directions

485 of the 1,110 sites run xSpeed Cache, which is ours, built by WPDeveloper. Both directions, under one caveat.

| Measure | xSpeed sites (485) | Everyone else (625) |
|:---|:---:|:---:|
| Median desktop score | **73.2** | 67.5 |
| Median mobile score | **58.0** | 54.6 |
| Median first byte | **274 ms** | 310 ms |
| Serving from cache | **83.3%** | 51.4% |
| Median mobile LCP | 7.3 s | **6.3 s** |
| Median Lighthouse gap | 17 pts | **15 pts** |

We score better on both views, answer faster and serve from cache far more often. We also have the **worse** median mobile LCP, by a second, and the **wider** Lighthouse gap. Sites install a caching plugin because they are slow, so this cohort was not drawn at random, and none of these figures separates the plugin from the reason it was installed.

The honest reading is the one the rest of this article argues for. Our cohort has largely finished the half of the problem the delivery rows measure, and is left with the half the device swap exposes.

## Running this across a fleet without opening forty dashboards

Reading two views on one report takes a minute. Reading them on forty client sites does not, which is the case [agencies](https://xspeedcache.com/use-cases/agencies/) run into first. [xSpeed Hub](https://xspeedcache.com/xspeed-hub/) holds every connected site behind one connection, and xSpeed Scan is also an MCP server documented at [Scan MCP](https://xspeedcache.com/docs/scan-mcp/), so an assistant can pull both graded views for a whole fleet and sort by the gap rather than by the headline. Both are ours.

Page caching and its exclusions are free and documented under [Page Cache](https://xspeedcache.com/docs/page-cache/); the module split is published on [Free vs Pro](https://xspeedcache.com/free-vs-pro/), and the full capability grid against eight competitors is on the [comparison page](https://xspeedcache.com/comparison/). For the delivery rows specifically, the honest answer is often the host rather than the plugin, and our standing recommendation there is [xCloud](https://xcloud.host/), which is also ours: xCloud and WPDeveloper are both Startise companies.

## Common mistakes people make with two-device reports

- 🖥️ **Acting on the headline.** It is the desktop run on all 1,129 reports that carry one, and Google indexes the other one.
- 📉 **Reading a small gap as good news.** The worst cohort in this study has the second-smallest gap.
- 🔁 **Re-testing until the columns agree.** They will not. They are different tests on different hardware.
- 🧮 **Treating a 4-point gap as a finding.** It sits inside the noise of a single run.
- 🖼️ **Assuming the desktop run is the lighter one.** It downloads more bytes on 9.9% of sites, because the viewport asks for bigger images.

## Frequently Asked Questions

### My desktop score is 88 and my mobile score is 59. Which one should I fix?

The mobile one, and the gap tells you where to start. A 29-point gap with the delivery rows identical in both views means the server is not the cause, so the work is JavaScript execution and image bytes rather than hosting.

### I assumed a big gap meant something was broken. Is it?

Not on its own. 60.3% of sites in this archive have a gap of 10 points or more, and the band with the widest gaps has the fastest servers. A wide gap mostly means the lab is punishing a page the desktop run flatters.

### I moved to a faster host and the gap did not change. Why not?

Because the gap cannot come from the host. The first-byte, cache, compression, redirect and CDN checks scored identically in both views on every site in the study. Faster hosting moves both columns together and leaves the difference between them alone.

### Does Google use the desktop score for anything?

Not for indexing. [Google Search Central](https://developers.google.com/search/docs/crawling-indexing/mobile/mobile-sites-mobile-first-indexing) states that the mobile version is what it crawls and ranks. The desktop run is useful for isolating a cause, which is exactly what this article uses it for.

### I re-ran the scan and both numbers changed. Did my site change?

Probably not. The same unchanged page returned 70, 61, 62 and 73 across four runs twelve minutes apart in [our benchmark protocol work](https://xspeedcache.com/blog/wordpress-caching-plugin-benchmark-protocol/). Treat any movement under about 12 points as the instrument.

### Why is my Cumulative Layout Shift worse on desktop?

It happens on 10.2% of sites. The desktop viewport lays the page out differently and can shift elements the narrow one does not. It is a real result about that layout, not a bug in the test.

### My page weight check passes on mobile and fails on desktop. Is that backwards?

It is correct and it is common, at 9.9% of sites. A 1350-pixel viewport picks larger `srcset` candidates than a 412-pixel one, so the desktop run really does transfer more. [Resizing beats format conversion](https://xspeedcache.com/blog/webp-images-page-still-heavy/) for the same reason.

### How many checks actually depend on the device?

Eleven of nineteen moved at least once. Eight never did. The Lighthouse performance score moved most, on 95.9% of sites, which follows from it being a weighted average of the metric scores, per [Chrome's Lighthouse scoring documentation](https://developer.chrome.com/docs/lighthouse/performance/performance-scoring).

### Is a cache plugin worth installing if my gap is wide?

It depends on the delivery rows, not the gap. If your first-byte and cache checks are already passing, a page cache has little left to take. We published [the headroom test](https://xspeedcache.com/blog/when-not-to-use-caching-plugin/) for exactly that question, and it says no for nearly half the archive.

### What threshold counts as passing?

Per [web.dev](https://web.dev/articles/lcp), Largest Contentful Paint should be 2.5 seconds or less, measured at *"the 75th percentile of page loads, segmented across mobile and desktop devices."* Google segments the threshold by device, which is the same reason a single blended score hides so much.

### Can I reproduce this on my own site?

Yes, and it costs nothing. Run the free [xSpeed Scan](https://xspeedcache.com/scan/), read the desktop and mobile views of the same report, and subtract. You can reproduce the underlying runs on Google's own PageSpeed Insights by passing each form factor.

## Conclusion: read the floor, then read the difference

A speed report with two device columns is not asking you to pick one. It is handing you a subtraction, and the subtraction separates the part of your score that follows the test hardware from the part that follows every visitor.

| If you want… | Do this | Why |
|:---|:---|:---|
| To know what to fix first | Read the mobile score, then the gap | The floor is the problem, the gap is the location |
| To rule out your host | Compare the delivery rows across the two views | They were identical on 100% of sites for five of six checks |
| To stop chasing noise | Ignore gaps under about 12 points | One run moves that much by itself |
| To do this for a fleet | Pull both views through the scan MCP | Forty dashboards is not a workflow |

**What to do this week:** run the free scan on your slowest page and read both views. Write down the mobile score and the difference. If the delivery rows match across the two views, stop looking at your host and start with the largest image and the heaviest script. If they do not match and both scores are low, the server is where the week goes. xSpeed Cache is free on [the WordPress.org directory](https://wordpress.org/plugins/xspeed/) and the paid tiers start at $29 a year with a 14-day money-back guarantee, on [our pricing page](https://xspeedcache.com/pricing/).

If you run this on your own archive and get a different answer, we would rather know. The method above is four numbers and a subtraction, and it is reproducible on any report the scan produces.
