research
How much traffic does a landing page test actually need?
Summary: At a 2–3% conversion rate you need roughly 300 landing page impressions per variant before a first-pass decision, and 5–10 conversions before you trust it. At $10k/day of paid spend — around 100,000 clicks — a 10% exploration budget supports about 30 variants in rotation. Below that traffic, the honest answer is that a classical two-variant A/B test on your landing page will not resolve, and most published CRO case studies quote effect sizes 3–10× larger than the real median.
Most advice about landing page test duration is "run it for two weeks" or "wait for significance." Neither is arithmetic. Here is the arithmetic.
Start with the effect size you are actually hunting
Sample size is a function of the effect you expect. Get this wrong and everything downstream is wrong.
The published median lift on a landing page A/B test is about 4%, and only 20–30% of tests reach statistical significance at all (Analytics-Toolkit's meta-analysis of 115 tests). Among tests that do win, lifts cluster at 5–15%, not the 25–35% that case studies advertise.
This matters more than any other number in this article. We maintain a calibrated set of 230 priors for landing page design choices — each an expected-lift estimate with a confidence rating and a source. The median expected lift is +1% across the ecommerce set and +2% across the service set. We cap any single prior at +20%, because when we built the first version from vendor case studies we ended up with entries like "video hero +30%" and "PAS headline +20%," and those numbers do not survive contact with experimental medians.
So: design your sample size for a 2–5% relative effect, not a 30% one. If you size for 30%, you will stop early, declare a winner, and be wrong.
The per-variant floor: why 300 impressions
Take a landing page converting at 2–3%. To make any judgement about a variant you need enough impressions that the conversion count is not pure noise.
300 impressions at a 2–3% conversion rate yields 6–9 conversions. That is the reasoning behind the 300-impression floor we use as a first-pass gate, treating one landing page impression as roughly equivalent to one ad click.
The related rule: aim for 5–10 conversions before making a decision on a variant. Not 5–10 conversions of margin — 5–10 conversions total. Below that, the posterior on that variant is so wide that "it is winning" and "it is losing" are both consistent with the data.
Here is why that threshold is not conservative enough to feel safe. Consider three arms with Beta-Bernoulli posteriors:
| Arm | Visits | Conversions | Posterior | Mean CVR |
|---|---|---|---|---|
| A | 50 | 5 | Beta(6, 46) | ~11.5% |
| B | 50 | 8 | Beta(9, 43) | ~17.3% |
| C | 10 | 2 | Beta(3, 9) | ~25% |
Arm C has the highest mean conversion rate in the table. Arm C is also the arm you know least about — its posterior is so wide it is nearly uninformative. A dashboard sorted by conversion rate puts C at the top. A team that ships C has learned nothing except that 2 out of 10 is a large fraction.
This is the single most common way landing page decisions go wrong, and it happens well above the 300-impression line.
The spend-to-variants calculation
Now scale up. Working numbers from paid social and display:
- $10,000/day of spend drives roughly 100,000 clicks on average.
- Reserve 10% for exploration — call it 9,000–10,000 impressions/day on unproven variants.
- At a 300-impression floor per variant, that exploration budget supports about 30 variants in rotation.
That is the whole calculation. Divide your daily exploration impressions by your per-variant floor, and you get your variant capacity.
Run it backwards and it becomes uncomfortable:
| Daily clicks | 10% explore | Variants supported at 300 imp |
|---|---|---|
| 100,000 | 10,000 | ~33 |
| 30,000 | 3,000 | ~10 |
| 10,000 | 1,000 | ~3 |
| 3,000 | 300 | 1 |
At 3,000 clicks/day you can put exactly one new variant into rotation per day and clear the floor. At 1,000 clicks/day you cannot clear it daily at all — you are accumulating for three days per variant before a first-pass look.
This is the real reason a classical A/B test fails on most landing pages. Not because the statistics are wrong, but because two variants a quarter, each needing weeks to accumulate a 2–5% detectable effect, is a learning rate of roughly nothing.
Why "statistically significant" arrives late and leaves early
Three structural problems compound the sample-size math, and none are fixed by waiting longer.
Conversions arrive late. A purchase can land seconds or days after the page view. Any decision rule reading conversions today is reading an incomplete label. Your 300 impressions have not finished converting.
Serving only the current best stops the data. If you route all traffic to the leader, you stop collecting on everything else. An unlucky-early variant never recovers, because it never gets the impressions that would have corrected the estimate. This is the failure mode of "declare a winner and ship it."
Performance drifts. Creative fatigues, audiences shift, seasons turn. A winner established over three weeks is a statement about those three weeks.
There is also a subtler trap. Effective sample size is not impression count. In sparse, high-dimensional corners — a specific creative × device × geo cell — importance weights blow up and effective sample size collapses even with millions of events. "We have millions of sessions" does not mean you have power in the cell you are trying to decide about.
The proxy metric that shortens the loop
If purchase conversions are too sparse to decide on, move up the funnel — carefully.
Click-through on the primary CTA fires for several-fold more visitors than purchase, so it accumulates fast enough that "is this page working?" stops being a three-week question. Its calibration is checkable in minutes rather than weeks.
The trap is well known and worth stating in numbers:
- Page A: 30% click × 10% buy = 3.0%
- Page B: 50% click × 4% buy = 2.0%
Clicks favour B. Revenue favours A. Optimising the proxy alone selects the clickbait page. So the proxy is for speed of learning, and the terminal metric still decides. A two-stage estimate — P(advance) then P(convert | advance) — gets you the accumulation rate of the proxy without letting it pick the winner.
What to do at each traffic level
Above ~50,000 clicks/day: you can support 15+ variants in exploration, promote the top ~10% into an exploitation pool, and retire anything that has passed 5,000 impressions while converting worse than control for three consecutive days. Hold a control group so you always have a live baseline.
5,000–50,000 clicks/day: you have a real but finite budget. Run 3–10 variants, accept that decisions take days not hours, and consider deciding on an upstream proxy with the purchase metric as a gate rather than the trigger.
Below ~5,000 clicks/day: a two-variant test targeting a 4% median effect will not resolve in any useful timeframe. Your honest options are to test changes big enough to produce a large effect (offer, mechanism, audience — not button colour), to pool learning across pages or brands rather than treating each test as independent, or to accept that you are making design decisions on priors and judgement, and stop performing statistics theatre about it.
That last option is more respectable than it sounds. A calibrated prior — even a median of +1–2% — is a real input. A test with 12% power is not.
The ramp we actually use
For reference, the allocation schedule we run when a campaign starts cold:
| Day 0 | Day 1 | Day 2 | Day 3 | Day 7 | Day 14 | |
|---|---|---|---|---|---|---|
| Control | 90% | 90% | 90% | 80% | 50% | 20% |
| Exploitation | 0% | 7% | 7% | 15% | 40% | 70% |
| Exploration | 10% | 3% | 3% | 5% | 10% | 10% |
Control stays dominant for three days because that is how long it takes to have anything worth exploiting. Exploration never goes to zero, because the moment it does you stop learning and start assuming.
Allocation is by Thompson sampling — randomise in proportion to how likely each arm is to be optimal, using Beta posteriors with an exponential decay so stale data fades. At a 0.95 daily decay factor, data from 14 days ago carries about 49% weight, which is how the drift problem gets handled without discarding history.
The short version
- Size for a 2–5% effect, because that is the real median. Sizing for 30% is why tests stop early and wrongly.
- 300 impressions per variant at a 2–3% CVR is a first-pass floor; 5–10 conversions is the decision threshold.
- Daily exploration impressions ÷ 300 = your variant capacity. At $10k/day, roughly 30.
- Below ~5,000 clicks/day, classical A/B testing on landing pages does not have the power to answer the question you are asking it.
- Millions of events does not mean power in the cell you care about.
Lutiq predicts, for every paid click, which on-brand page to show — and measures lift against a live holdout. Figures above are first-party unless attributed.