lutiq

research

We rebuilt 230 CRO benchmarks against the research. The median real effect is +1–2%.

· research

Summary: We needed expected-lift estimates for every landing page design choice our generator can make. Version one was built from published case studies and produced numbers like "video hero +30%" and "personalised CTA +25%." When we recalibrated against experimental medians and meta-analyses, we had to cut almost every estimate by 3–10× and impose a hard +20% ceiling. Across the resulting 230 priors, the median expected lift is +1% for ecommerce and +2% for service. This is a quantified account of how far CRO case-study numbers sit from CRO reality — including our own first attempt.

Any system that generates landing pages needs an opinion about what works before it has data. We encode those opinions as priors: for each design dimension and variation, an expected lift, a confidence rating, and a source.

Building that set twice taught us more about the CRO literature than reading it once did.

Version one: what happens when you believe the case studies

The first prior set expanded from 12 entries across 6 dimensions to about 134 entries across 18 dimensions, sourced the obvious way — from published CRO case studies and vendor blog posts.

Representative entries:

Variationv1 expected liftSource shape
Video hero+30%vendor case study
Personalised CTA+25%"+42% higher CVR, HubSpot"
PAS headline+20%">400% more CTA clicks"
Sticky CTA+20%"+37% checkout starts"

Every one of those has a citation. None of them are fabricated. And the set as a whole was badly wrong, because of how case studies are selected for publication: a vendor publishes the test that produced +42%, not the eleven that produced +1%, −3%, and no significant difference. Aggregate a corpus of published wins and you get a corpus of outliers.

What the experimental record actually says

The recalibration anchor was meta-analysis rather than case study:

Against that, a prior set with a +30% median is not optimistic. It is a different distribution entirely.

Version two: the recalibration

We rebuilt the taxonomy from 19 sprawling dimensions and ~156 priors down to 12 orthogonal dimensions and ~95 priors for ecommerce, plus a parallel service set, each entry re-sourced and re-tiered. Two rules did most of the work:

A hard +20% ceiling on any single prior. No design choice gets credit for more than a 20% expected lift, regardless of what any case study claims. A change that genuinely produces more than 20% is almost always an offer or mechanism change, not a design change — and it should be tested as such rather than smuggled in as "hero treatment."

Source tiering with explicit confidence. Every prior carries high, medium, or low confidence based on whether the effect replicates across independent sources.

The resulting distribution across 230 priors (114 ecommerce, 116 service):

EcommerceService
Priors114116
Dimensions1919
Families4849
Expected lift range−0.10 to +0.15−0.15 to +0.18
Median expected lift+2%+2%
High confidence3433
Medium confidence4652
Low confidence3431

The median is +1% for the ecommerce set and +2% for service. Only about 29% of entries are high confidence. And a meaningful number are negative — design choices we expect to hurt, which a case-study-derived set will essentially never contain, because nobody publishes "we added seven trust badges and conversion fell."

Effects that survived recalibration

Not everything got cut. These replicated across independent sources and kept high confidence — note that the ones that survive are mostly about audience and context, not aesthetics:

The mobile gap. Mobile is ≥73% of ecommerce traffic but converts at roughly half the desktop rate — 1.53% mobile vs 4.31% desktop (ContentSquare 2025, 90B sessions across 6,000 sites). This is the largest single structural opportunity in the data, and it is not a design pattern.

Traffic temperature. Email and branded traffic converts about 4× higher than cold paid traffic (landing-page platform benchmark data, 2024). Any benchmark quoted without a traffic source is close to meaningless.

Reading level. Copy at 5th–7th grade reading level converts about 2× better than college-level; in SaaS specifically the gap is 12.9% vs 2.1% — a 6× difference on that dataset.

Above-fold attention. Above-fold content receives 84% of viewing time; headlines are read more than 5× as often as body copy.

Trust signal saturation. 1–3 trust signal types is optimal; 7+ types reduces conversion (Baymard). Non-monotonic, which is why "add more badges" fails.

Form length. 3-field forms convert at 25% against 9-field forms at 3.6%; each field beyond five costs 4–7%.

CTA focus. Single-CTA pages convert at 13.5% vs 10.5% for pages with 5+ CTAs.

Above-fold reviews. +164% lift on click (PowerReviews 2024) — a click metric, not a purchase metric, and the distinction matters.

Where the evidence is genuinely weak

The more useful output of the recalibration was cataloguing what nobody has actually established. Three areas where the confident advice outruns the data:

Thumb-zone placement is running on 2013–2014 research. The foundational studies predate Dynamic Island, the Action Button, and the dominance of in-app browsers. The finding that a sticky bottom CTA lifts conversion on long pages by +17–27% does replicate. The claim about optimal position within the thumb zone does not — it rests on device ergonomics that no longer describe the devices people use.

"Authentic photography beats stock by +35%" may be decaying. That effect was measured before generative imagery was good. No rigorous 2024–2026 replication exists, and the mechanism — perceived authenticity — is exactly what improving image generation erodes.

B2B device splits are shifting under the research. The 65/35 desktop-mobile assumption comes from 2024 data and is likely moving as younger buyers enter procurement.

And three specific numbers that circulate widely and should not:

The countdown timer "+332%" case. Real test, single product, unusual context. The median for urgency mechanics is +5–15%. The 332% figure is the single most misleading number in circulation in CRO.

B2B chat conversion claims of 10–25%. These conflate visitors who engaged with chat against page-level conversion rate. Different denominators. The comparison is invalid as usually presented.

Quiz lift figures of +30–47%. Vendor-self-reported, from companies selling quiz software. We kept quizzes as a medium-confidence positive prior, not high.

Why this matters if you are not building a generator

Three practical consequences.

Your test is probably underpowered because your expected effect is wrong. Size a test for +30% and you need a fraction of the sample you actually need for +4%. You will hit "significance" early on noise, ship, and regress. Sizing for the real median is the single highest-leverage change most teams could make to their testing programme.

A pattern with a big published lift is not a pattern with a big expected lift. The published number is conditioned on having been published. Assume regression toward +2–5% and treat anything above +20% as a claim about an offer, not a design.

Confidence should be carried alongside every benchmark. We rate 34 of 114 ecommerce priors as high confidence. Any benchmark list that presents all of its entries with equal authority is hiding the thing you most need to know.

Method and limitations

The prior set is machine-readable — each entry is {dimension, family, variation, expected_lift, confidence, reasoning} with the reasoning naming its source. Priors are starting points that get overwritten by a brand's own data as it accumulates; a prior that survives contact with 20,000 visitors is no longer a prior.

Two limitations worth stating. These are calibrated estimates, not measurements — our contribution is the synthesis, the tiering, and the ceiling, not new experiments behind each number. And the underlying corpus skews toward companies that publish CRO research, which skews ecommerce, English-language, and higher-traffic. The mobile-gap and traffic-temperature findings are the most robust because they come from the largest and least self-selected datasets.


Lutiq predicts, for every paid click, which on-brand page to show — and measures lift against a live holdout. Figures above are first-party unless attributed.