How Redline measures a device
This page documents exactly what the Redline lab runs when you press “Stress this device”, how the numbers become a score, and where that score should not be trusted. It describes the code that is live on this site today.
By Redline Labs · Published
Principles
- Everything in the device benchmark runs in your browser tab, on the main JavaScript thread, using standard web APIs. Nothing is installed.
- Each test is time-boxed rather than work-boxed: it does as much work as it can in a fixed window and reports a rate. Slow devices finish in the same time; they just get less done.
- Raw rates are kept alongside scores, so you can compare devices on the underlying numbers rather than only the 0–100 scale.
- The benchmark measures device capacity, not your site. The separate URL probe looks at how a specific page is delivered.
What is read before the tests
Before running, Redline reads what the browser is willing to report: logical core count (navigator.hardwareConcurrency), approximate memory (navigator.deviceMemory, Chromium only, rounded and capped by the browser), the effective connection type and downlink estimate where the Network Information API exists, screen size and pixel ratio, touch support, and the WebGL renderer string if the browser exposes it. Form factor is inferred from touch support and screen width (touch and under 768 px is treated as a phone, under 1,180 px as a tablet, anything else as desktop). Many of these values are deliberately coarse or missing in Safari and Firefox for privacy reasons; Redline shows “unknown” rather than guessing.
The five tests
1. CPU throughput
A tight loop of mixed floating-point and integer work (a square root, a modulo and a bitwise XOR per iteration) runs in batches of 20,000 operations for about 900 ms. The result is millions of operations per second. Because it runs on the main thread, it reflects single-core speed plus whatever the browser’s JIT compiler does with the loop — the same constraints your own JavaScript runs under. It does not use Web Workers, so it does not measure multi-core throughput.
2. Memory bandwidth
Up to 24 blocks of one million 64-bit floats (8 MB each, so up to 192 MB) are allocated, and every 64th element is written, which touches each cache line. The test stops early after about 900 ms. The score is megabytes allocated and touched per second. A low score usually means slow memory, aggressive garbage collection, or a browser that is close to its per-tab memory limit — all of which also hurt image-heavy and data-heavy pages.
3. DOM / layout
In an off-screen container, Redline repeatedly appends 200 transformed span elements and then reads offsetHeight, which forces the browser to recalculate layout synchronously. This loop runs for about 700 ms and reports nodes created and laid out per second. It imitates the most common cause of jank on real sites: code that writes to the DOM and then immediately reads geometry.
4. Graphics / frame rate
A 640×360 canvas receives 700 coloured rectangle fills per frame for one second, paced by requestAnimationFrame. Redline records every frame duration and reports average frames per second plus the single worst frame. The fps score is capped by the display’s refresh rate on most browsers, so a 120 Hz device can reach the 60 fps reference easily; the worst-frame figure is often the more revealing number.
5. Network latency
Six small requests are sent to Redline’s own origin with caching disabled, 40 ms apart. The median round trip is reported, along with jitter (slowest minus fastest). This is a latency check against one server, not a bandwidth test and not a measurement of your site’s server.
How the composite score is built
Each raw rate is divided by a fixed reference value and capped to 0–100. Network uses the formula 100 − (median ms − 20) ÷ 3, so roughly 20 ms scores 100 and 320 ms scores 0. The five scores are then combined with fixed weights:
| Test | Time budget | Reference | Weight |
|---|---|---|---|
| CPU throughput | ≈ 0.9 s | 60 Mops/s = 100 | 28% |
| DOM / layout | ≈ 0.7 s | 90,000 nodes/s = 100 | 24% |
| Graphics / frame rate | ≈ 1.0 s | 60 fps = 100 | 22% |
| Memory bandwidth | ≤ 0.9 s | 4,000 MB/s = 100 | 16% |
| Network latency | 6 probes | 20 ms median ≈ 100 | 10% |
The weighted total is rounded and mapped to a grade: S at 88 or above, A at 75, B at 60, C at 42, and D below 42. The weights favour CPU, layout and graphics because those are the resources most often exhausted by modern JavaScript-heavy pages; network carries less weight because a single latency probe says little about how a given site will load.
What the score is not
Variance and repeatability
Short, time-boxed browser benchmarks are noisy. In practice, expect a few points of movement between consecutive runs on the same device, and more on phones. The main sources are:
- Thermal state: phones throttle under sustained load, so the third run in a row is often lower than the first.
- Power mode: battery saver and low-power modes cap CPU frequency and frame rate.
- Background work: other tabs, extensions, sync and OS updates compete for the same cores.
- JIT warm-up and garbage collection timing, which differ from run to run.
- Network conditions for the latency test, which change minute to minute on mobile.
Recommended test conditions
- Close other tabs and apps; keep the Redline tab in the foreground for the whole run.
- Plug in or disable battery saver, unless low-power behaviour is what you want to measure.
- Let the device cool for a minute between runs if you are comparing devices.
- Run three times and use the median overall score; discard a run if a notification or app switch interrupted it.
- Record the browser and version — Chrome, Safari and Firefox on the same hardware can differ noticeably on the layout and graphics tests.
The URL probe is a different measurement
The URL probe does not run on your device. It sends the address you enter to the Redline edge, which fetches the page once and grades nine delivery checks: time to first byte, document transfer size, compression, render-blocking scripts, stylesheets in the head, lazy-loaded images, third-party origins, cache policy and inline CSS. It reads the HTML document only; it does not execute the page’s JavaScript or load its images, so it is a delivery check, not a full page-load audit. The public CI endpoint applies the same checks against budgets you set.
How this differs from Lighthouse and load testing
Lighthouse loads one page in one controlled browser and simulates a slower device and network to produce lab metrics for that page. Load-testing tools such as k6 and JMeter generate many concurrent requests to find where a server slows down. Redline measures the opposite end: how much compute, memory, layout and graphics capacity the actual device in someone’s hand has. The three answer different questions and are best used together. Our guides on load testing versus real-device testing and Core Web Vitals go into more detail.
Limitations — when not to use Redline
- Not for server capacity planning — it sends no meaningful load to any server.
- Not a substitute for field data: real-user Core Web Vitals from your own analytics or the Chrome UX Report remain the ground truth.
- Not a hardware certification or a cross-browser engine comparison; browser differences are part of the result.
- Not a multi-core or GPU-shader benchmark: CPU runs on one thread and graphics uses 2D canvas fills.
- Not reliable in a background tab — browsers throttle timers and animation frames there.
- Device details such as memory and GPU name are whatever the browser reports, and may be rounded, hidden or spoofed.
Reproducing and interpreting a run
To reproduce a result, run the lab on the same device and browser under the conditions above, three times, and compare the median overall score and the raw rate for each test. When interpreting a run, look first at the lowest individual score rather than the composite: a device with a strong CPU but a weak layout score will struggle with DOM-heavy pages even if its overall grade looks healthy. The guide to reading a device performance score walks through worked examples.
The optional written analysis on paid plans is generated by an AI model from the scores and the device profile. It is an interpretation aid, can be wrong, and is never used to compute the score.
Who maintains this page
This documentation is written by Redline Labs, the team that builds the benchmark. When the test code, reference values or weights change, this page is updated in the same release and its updated date changes. If you find a discrepancy, please tell us.