Redline runs the same stress on the actual hardware your users carry. It shows you exactly which subsystem will drop frames on a low-end phone, and which fix pays off there — instead of the fix that only felt faster on your machine.
Run a stress testYour dev machine is a best case. Run Redline on the cheapest phone you support and set CI budgets against that score, not against an M2 MacBook.
The CPU benchmark pushes mixed math until the thread blocks, so you can watch input latency emerge on real silicon — not guess at it from a flame chart.
The DOM/layout benchmark forces reflow on the device's own engine. If a deep component tree drops frames here, it drops frames in production.
The analysis references your real numbers and names concrete techniques — scheduler.yield(), contain, transform-only animations — so you can act immediately.
A current developer laptop has somewhere between eight and sixteen fast cores, tens of gigabytes of memory, a discrete or high-end integrated GPU, an always-warm cache and a wired or strong Wi-Fi connection. A mid-range Android sold in volume has four efficiency cores and two performance cores that throttle within a minute, four gigabytes of shared memory, a mobile GPU an order of magnitude slower, and a network that varies by the minute. Single-thread JavaScript execution alone commonly differs by a factor of six to ten.
That gap is not a rounding error — it changes which optimisation is correct. A hydration cost of 40 ms on your laptop is 300 ms on that phone, which is the difference between an invisible delay and a tap that appears to do nothing. Measuring on the floor rather than the ceiling is the entire reason Redline exists.
Step 1
Pick the floor device
Open your analytics, sort device models by sessions, and take the slowest model in the bottom quartile of your traffic. That device, not the newest flagship, is the one your budget is written against.
Step 2
Record a baseline
Run the Redline stress test on that device and save the run. You now have CPU, memory, layout, GPU and network sub-scores, plus an overall index, from before any change.
Step 3
Change one thing
Ship a single change — defer a third-party script, split a bundle, replace a layout-triggering animation with a transform. One change at a time keeps the attribution honest.
Step 4
Re-run and compare
Run the same test on the same device, ideally at a similar battery level and temperature. Compare sub-scores, not just the headline index: a change can help layout and hurt memory.
Step 5
Gate it
Once a number is trustworthy, POST it as a budget to the Redline audit endpoint from CI so a regression fails the build instead of reaching a user.
A team animated a drawer by transitioning height and top. On a desktop the animation was a smooth 60 fps and nobody questioned it. On a mid-tier phone the layout sub-score dropped by roughly a third during the interaction, because every frame re-ran layout for the whole subtree beneath the drawer.
Rewriting the same visual effect with transform: translateY() and opacity, and adding contain: layout paint to the drawer, restored the layout sub-score and removed the dropped frames. The visible design did not change at all. Nothing in a desktop profile would have prompted the work — the device measurement did.
Redline measures device capability under a controlled synthetic workload. It does not replay your application's own code paths, so it cannot tell you that a particular React component re-renders too often, and it does not observe real users in the field. Pair it with a profiler for component-level attribution and with field data (RUM or the Chrome UX Report) for what users actually experience.
Scores also vary between runs — thermal state, battery saver, other tabs and background apps all move the number. Treat a single run as an estimate and a run-to-run difference under roughly five points as noise.
What the CPU, memory, layout, GPU and network numbers in a device benchmark actually mean, and how to tell a real problem from normal variance.
The hardware gap, thermal throttling, network variance and cache warmth that make developer machines a misleading place to judge performance.
An ordered checklist for making a page load faster, with the realistic gain and the cost of each step so you can stop at the right point.
LCP, INP and CLS explained without jargon: what each one measures, what usually breaks it, and the change that most often moves it.