REDLINE
For QA engineers

A performance regression you can actually point at.

Load testing already covers your server. Redline covers the other half: the device. Run a repeatable stress test on real hardware, keep the baseline, and turn “it feels slower” into a number you can ticket.

Run a baseline test

A baseline you can reproduce

Run the same five-subsystem stress on any device and record the Redline Index. Every future run compares against that baseline, so regressions show up as a number — not a vibe.

Catches what load testing misses

Load tools tell you the server held up. Redline tells you the cheap phone didn't. Together they cover both ends of the delivery path.

Shareable reports, no setup

Send a stakeholder a Redline report instead of a screenshot of a flame chart. The verdict, cohort and fixes are written in plain engineering language.

CI budgets, shipping today

POST a URL and a performance budget to the Redline audit endpoint. It returns 412 when the budget is breached, so a regression fails the build instead of reaching production.

Gate a release on a budget

Add this to any pipeline. A non-zero exit means the delivery budget was breached.

curl -sS -X POST https://redline.lovable.app/api/public/v1/audit \
  -H 'content-type: application/json' \
  -d '{"url":"https://example.com","budget":{"minScore":80,"maxTtfbMs":600}}' \
  -o report.json -w '%{http_code}' | grep -q 200

Writing a device test that is actually repeatable

Performance tests get abandoned because their results wander. Most of that wander is controllable. Before a run, put the device on mains power or a consistent battery level, close other applications and tabs, disable battery saver, and let the device sit for a minute so it is not still hot from the previous run. Use the same browser and the same network each time.

Then run three times and record the median, never the best. A single run is an observation; three runs is a measurement. If the spread between the highest and lowest run is wider than about ten points, something in the environment is uncontrolled — find it before you trust the number.

How much movement is a real regression?

As a working rule on a stable setup: under five points is noise, five to ten points is worth a second look, and more than ten points on the same device with the same build settings is a regression to file. Always compare sub-scores too — an overall index can stay flat while the layout score collapses and the network score improves.

A device matrix that does not need a device lab

You do not need fifty handsets. Four tiers cover almost every realistic decision: a current flagship, a two-to-three-year-old mid-range phone, the cheapest device still present in your analytics, and a tablet or low-end laptop if your product sees desktop traffic. Record a baseline for each, refresh it every release train, and keep the histories side by side.

The floor device is the one that sets the budget. The flagship exists only to prove that a regression is device-specific rather than universal — when both drop, the cause is in the code path; when only the floor device drops, the cause is a resource or capability limit.

Turning a run into a ticket

A useful performance ticket contains the device model and OS version, the browser and version, the build or commit under test, the baseline index and sub-scores, the current index and sub-scores, the median of three runs, and the shared Redline report link. With those fields a developer can reproduce your result; without them the ticket becomes a debate.

Include the plain-language verdict from the report as well. It gives non-engineering reviewers enough context to prioritise the work without asking someone to translate a flame chart.

Honest limitations

Redline runs inside a browser, so it measures what a browser can measure. It cannot see native app behaviour, it cannot instrument your server, and it cannot simulate a thousand concurrent users — that is what load tools are for. It also cannot prove a fix works for real users; only field data can do that. Use it as the device-side half of a performance practice, not the whole of one.

Read next