REDLINE
For frontend developers

Your laptop is lying to you about performance.

Redline runs the same stress on the actual hardware your users carry. It shows you exactly which subsystem will drop frames on a low-end phone, and which fix pays off there — instead of the fix that only felt faster on your machine.

Run a stress test

Set budgets from the floor, not the ceiling

Your dev machine is a best case. Run Redline on the cheapest phone you support and set CI budgets against that score, not against an M2 MacBook.

See the main thread actually saturate

The CPU benchmark pushes mixed math until the thread blocks, so you can watch input latency emerge on real silicon — not guess at it from a flame chart.

Catch layout thrash before it ships

The DOM/layout benchmark forces reflow on the device's own engine. If a deep component tree drops frames here, it drops frames in production.

AI fixes that name the API, not the vibe

The analysis references your real numbers and names concrete techniques — scheduler.yield(), contain, transform-only animations — so you can act immediately.

The gap between your machine and your user's machine

A current developer laptop has somewhere between eight and sixteen fast cores, tens of gigabytes of memory, a discrete or high-end integrated GPU, an always-warm cache and a wired or strong Wi-Fi connection. A mid-range Android sold in volume has four efficiency cores and two performance cores that throttle within a minute, four gigabytes of shared memory, a mobile GPU an order of magnitude slower, and a network that varies by the minute. Single-thread JavaScript execution alone commonly differs by a factor of six to ten.

That gap is not a rounding error — it changes which optimisation is correct. A hydration cost of 40 ms on your laptop is 300 ms on that phone, which is the difference between an invisible delay and a tap that appears to do nothing. Measuring on the floor rather than the ceiling is the entire reason Redline exists.

A workflow that fits an ordinary sprint

  1. Step 1

    Pick the floor device

    Open your analytics, sort device models by sessions, and take the slowest model in the bottom quartile of your traffic. That device, not the newest flagship, is the one your budget is written against.

  2. Step 2

    Record a baseline

    Run the Redline stress test on that device and save the run. You now have CPU, memory, layout, GPU and network sub-scores, plus an overall index, from before any change.

  3. Step 3

    Change one thing

    Ship a single change — defer a third-party script, split a bundle, replace a layout-triggering animation with a transform. One change at a time keeps the attribution honest.

  4. Step 4

    Re-run and compare

    Run the same test on the same device, ideally at a similar battery level and temperature. Compare sub-scores, not just the headline index: a change can help layout and hurt memory.

  5. Step 5

    Gate it

    Once a number is trustworthy, POST it as a budget to the Redline audit endpoint from CI so a regression fails the build instead of reaching a user.

A worked example: the 'fast' animation that wasn't

A team animated a drawer by transitioning height and top. On a desktop the animation was a smooth 60 fps and nobody questioned it. On a mid-tier phone the layout sub-score dropped by roughly a third during the interaction, because every frame re-ran layout for the whole subtree beneath the drawer.

Rewriting the same visual effect with transform: translateY() and opacity, and adding contain: layout paint to the drawer, restored the layout sub-score and removed the dropped frames. The visible design did not change at all. Nothing in a desktop profile would have prompted the work — the device measurement did.

What Redline does not tell you

Redline measures device capability under a controlled synthetic workload. It does not replay your application's own code paths, so it cannot tell you that a particular React component re-renders too often, and it does not observe real users in the field. Pair it with a profiler for component-level attribution and with field data (RUM or the Chrome UX Report) for what users actually experience.

Scores also vary between runs — thermal state, battery saver, other tabs and background apps all move the number. Treat a single run as an estimate and a run-to-run difference under roughly five points as noise.

Read next