research
7,768 impressions, 52 conversions: why landing-page allocator backtests go quiet
Summary: Multi-variant landing page systems live in a conversion desert. In one corrected walk-forward evaluation of production-style allocation logs we retained 7,768 eligible impressions and only 52 primary-endpoint conversions. Apparent lifts of +19% or even +47% on small confirmation slices routinely fail bootstrap uncertainty tests. This piece is about power and honest methodology — not a claim that a particular model “won.”
Paid landing pages convert at a few percent. Multi-page races split that further. By the time you enforce mature labels, reconstructable propensities, multi-page decision windows and daily walk-forward eligibility, the conversion count collapses.
That is not a tooling problem. It is the measurement problem.
What we measured
We rebuilt an impression table from production-style allocation logs (anonymised multi-experiment sample): on the order of ~10,750 raw production impressions across 61 pages and 10 experiments, with a small absolute conversion count (tens of purchases/leads inside a 72-hour window).
After label maturity, propensity reconstruction, multi-page decision windows and daily-window eligibility, the full 24-hour evaluation retained:
| Quantity | Value |
|---|---|
| Eligible impressions | 7,768 |
| Primary-endpoint conversions | 52 |
| Pages in race pool (raw table) | 61 |
| Experiments | 10 |
Open data summary: allocator-power-2026-07.json.
Method safeguards that change the story
Several “optimistic” mistakes turn noise into a press release:
- Treating immature non-conversions as negatives before the full attribution window closes.
- Mis-reading partial history rows as complete allocation states.
- Picking the policy on the same slice you report (selection leakage).
- Ignoring that production history may not reconstruct arm shares after re-wires.
After corrections, training is walk-forward at a fixed daily boundary; non-conversions only enter training after a complete attribution window; unknown propensities are excluded rather than guessed.
Estimators used SNIPS against logged allocation probabilities, with bootstrap uncertainty by experiment × day (visitor bootstrap as sensitivity).
What “significant” looked like after correction
Representative corrected results (not a product claim):
| Slice / policy style | Point lift | Uncertainty signature |
|---|---|---|
| Later confirmation, greedy | +47.1% | Only 21 conversions on that slice; bootstrap P(no uplift) ~ 0.13; CI crosses zero |
| Same family, soft policy on later slice | −2.6% | Inverts the earlier selection story |
| Full usable window, soft | +2.6% | Not significant |
| Full usable window, greedy | +19.0% | Not significant |
A doubly robust sensitivity analysis moved some point estimates but did not produce stable, significant uplift under visitor-level uncertainty.
Interpretation
- Conversion rarity dominates model choice. With ~52 primary conversions, most policy comparisons are underpowered. Publishing a single +X% without n and uncertainty is marketing, not Research.
- Temporal instability is the default. Selection-slice winners that invert on confirmation are a methodology failure mode, not “the model needs more layers.”
- Holdouts still matter. Off-policy estimators on thin multi-page data can be confidently wrong. A live holdout is the backstop — see the Lutiq Holdout Protocol.
- Related sample-size arithmetic for classical tests lives in How much traffic does a landing page test actually need?.
Caveats
- Anonymised multi-experiment production-style logs; not a public raw dump.
- Primary endpoint is conversion within a fixed attribution window, not revenue-optimised.
- Numbers characterise power and methodology, not a guaranteed allocator ROI.