REDLINE
Reference documentation

How Redline measures a device

This page documents exactly what the Redline lab runs when you press “Stress this device”, how the numbers become a score, and where that score should not be trusted. It describes the code that is live on this site today.

By Redline Labs · Published

Principles

What is read before the tests

Before running, Redline reads what the browser is willing to report: logical core count (navigator.hardwareConcurrency), approximate memory (navigator.deviceMemory, Chromium only, rounded and capped by the browser), the effective connection type and downlink estimate where the Network Information API exists, screen size and pixel ratio, touch support, and the WebGL renderer string if the browser exposes it. Form factor is inferred from touch support and screen width (touch and under 768 px is treated as a phone, under 1,180 px as a tablet, anything else as desktop). Many of these values are deliberately coarse or missing in Safari and Firefox for privacy reasons; Redline shows “unknown” rather than guessing.

The five tests

1. CPU throughput

A tight loop of mixed floating-point and integer work (a square root, a modulo and a bitwise XOR per iteration) runs in batches of 20,000 operations for about 900 ms. The result is millions of operations per second. Because it runs on the main thread, it reflects single-core speed plus whatever the browser’s JIT compiler does with the loop — the same constraints your own JavaScript runs under. It does not use Web Workers, so it does not measure multi-core throughput.

2. Memory bandwidth

Up to 24 blocks of one million 64-bit floats (8 MB each, so up to 192 MB) are allocated, and every 64th element is written, which touches each cache line. The test stops early after about 900 ms. The score is megabytes allocated and touched per second. A low score usually means slow memory, aggressive garbage collection, or a browser that is close to its per-tab memory limit — all of which also hurt image-heavy and data-heavy pages.

3. DOM / layout

In an off-screen container, Redline repeatedly appends 200 transformed span elements and then reads offsetHeight, which forces the browser to recalculate layout synchronously. This loop runs for about 700 ms and reports nodes created and laid out per second. It imitates the most common cause of jank on real sites: code that writes to the DOM and then immediately reads geometry.

4. Graphics / frame rate

A 640×360 canvas receives 700 coloured rectangle fills per frame for one second, paced by requestAnimationFrame. Redline records every frame duration and reports average frames per second plus the single worst frame. The fps score is capped by the display’s refresh rate on most browsers, so a 120 Hz device can reach the 60 fps reference easily; the worst-frame figure is often the more revealing number.

5. Network latency

Six small requests are sent to Redline’s own origin with caching disabled, 40 ms apart. The median round trip is reported, along with jitter (slowest minus fastest). This is a latency check against one server, not a bandwidth test and not a measurement of your site’s server.

How the composite score is built

Each raw rate is divided by a fixed reference value and capped to 0–100. Network uses the formula 100 − (median ms − 20) ÷ 3, so roughly 20 ms scores 100 and 320 ms scores 0. The five scores are then combined with fixed weights:

TestTime budgetReferenceWeight
CPU throughput≈ 0.9 s60 Mops/s = 10028%
DOM / layout≈ 0.7 s90,000 nodes/s = 10024%
Graphics / frame rate≈ 1.0 s60 fps = 10022%
Memory bandwidth≤ 0.9 s4,000 MB/s = 10016%
Network latency6 probes20 ms median ≈ 10010%

The weighted total is rounded and mapped to a grade: S at 88 or above, A at 75, B at 60, C at 42, and D below 42. The weights favour CPU, layout and graphics because those are the resources most often exhausted by modern JavaScript-heavy pages; network carries less weight because a single latency probe says little about how a given site will load.

What the score is not

The references are fixed engineering targets, not percentiles and not a calibration against a device database. Scores are capped, so two fast desktops can both read 100 while differing substantially in raw throughput. A score is a relative indicator of headroom, not a prediction of a Lighthouse or Core Web Vitals result. Where enough runs have been saved to Redline, the results page can also show how a run compares with other saved runs; that comparison reflects only who chose to save runs, not the device market as a whole.

Variance and repeatability

Short, time-boxed browser benchmarks are noisy. In practice, expect a few points of movement between consecutive runs on the same device, and more on phones. The main sources are:

Recommended test conditions

The URL probe is a different measurement

The URL probe does not run on your device. It sends the address you enter to the Redline edge, which fetches the page once and grades nine delivery checks: time to first byte, document transfer size, compression, render-blocking scripts, stylesheets in the head, lazy-loaded images, third-party origins, cache policy and inline CSS. It reads the HTML document only; it does not execute the page’s JavaScript or load its images, so it is a delivery check, not a full page-load audit. The public CI endpoint applies the same checks against budgets you set.

How this differs from Lighthouse and load testing

Lighthouse loads one page in one controlled browser and simulates a slower device and network to produce lab metrics for that page. Load-testing tools such as k6 and JMeter generate many concurrent requests to find where a server slows down. Redline measures the opposite end: how much compute, memory, layout and graphics capacity the actual device in someone’s hand has. The three answer different questions and are best used together. Our guides on load testing versus real-device testing and Core Web Vitals go into more detail.

Limitations — when not to use Redline

Reproducing and interpreting a run

To reproduce a result, run the lab on the same device and browser under the conditions above, three times, and compare the median overall score and the raw rate for each test. When interpreting a run, look first at the lowest individual score rather than the composite: a device with a strong CPU but a weak layout score will struggle with DOM-heavy pages even if its overall grade looks healthy. The guide to reading a device performance score walks through worked examples.

The optional written analysis on paid plans is generated by an AI model from the scores and the device profile. It is an interpretation aid, can be wrong, and is never used to compute the score.

Who maintains this page

This documentation is written by Redline Labs, the team that builds the benchmark. When the test code, reference values or weights change, this page is updated in the same release and its updated date changes. If you find a discrepancy, please tell us.

Run the benchmarkRead the guides