lutiq

research

7,768 impressions, 52 conversions: why landing-page allocator backtests go quiet

· research

Summary: Multi-variant landing page systems live in a conversion desert. In one corrected walk-forward evaluation of production-style allocation logs we retained 7,768 eligible impressions and only 52 primary-endpoint conversions. Apparent lifts of +19% or even +47% on small confirmation slices routinely fail bootstrap uncertainty tests. This piece is about power and honest methodology — not a claim that a particular model “won.”

Paid landing pages convert at a few percent. Multi-page races split that further. By the time you enforce mature labels, reconstructable propensities, multi-page decision windows and daily walk-forward eligibility, the conversion count collapses.

That is not a tooling problem. It is the measurement problem.

What we measured

We rebuilt an impression table from production-style allocation logs (anonymised multi-experiment sample): on the order of ~10,750 raw production impressions across 61 pages and 10 experiments, with a small absolute conversion count (tens of purchases/leads inside a 72-hour window).

After label maturity, propensity reconstruction, multi-page decision windows and daily-window eligibility, the full 24-hour evaluation retained:

QuantityValue
Eligible impressions7,768
Primary-endpoint conversions52
Pages in race pool (raw table)61
Experiments10

Open data summary: allocator-power-2026-07.json.

Method safeguards that change the story

Several “optimistic” mistakes turn noise into a press release:

  1. Treating immature non-conversions as negatives before the full attribution window closes.
  2. Mis-reading partial history rows as complete allocation states.
  3. Picking the policy on the same slice you report (selection leakage).
  4. Ignoring that production history may not reconstruct arm shares after re-wires.

After corrections, training is walk-forward at a fixed daily boundary; non-conversions only enter training after a complete attribution window; unknown propensities are excluded rather than guessed.

Estimators used SNIPS against logged allocation probabilities, with bootstrap uncertainty by experiment × day (visitor bootstrap as sensitivity).

What “significant” looked like after correction

Representative corrected results (not a product claim):

Slice / policy stylePoint liftUncertainty signature
Later confirmation, greedy+47.1%Only 21 conversions on that slice; bootstrap P(no uplift) ~ 0.13; CI crosses zero
Same family, soft policy on later slice−2.6%Inverts the earlier selection story
Full usable window, soft+2.6%Not significant
Full usable window, greedy+19.0%Not significant

A doubly robust sensitivity analysis moved some point estimates but did not produce stable, significant uplift under visitor-level uncertainty.

Interpretation

  1. Conversion rarity dominates model choice. With ~52 primary conversions, most policy comparisons are underpowered. Publishing a single +X% without n and uncertainty is marketing, not Research.
  1. Temporal instability is the default. Selection-slice winners that invert on confirmation are a methodology failure mode, not “the model needs more layers.”
  1. Holdouts still matter. Off-policy estimators on thin multi-page data can be confidently wrong. A live holdout is the backstop — see the Lutiq Holdout Protocol.
  1. Related sample-size arithmetic for classical tests lives in How much traffic does a landing page test actually need?.

Caveats