lutiq

glossary

What is a holdout test (for landing pages)?

· glossary

A holdout test reserves a portion of traffic for a baseline experience while the rest of traffic receives the new treatment (a new page, a new router, or a multi-variant system). You compare outcomes between holdout and treatment to estimate true lift.

Without a holdout, teams often declare victory because variant B beat variant A inside the experiment — while the whole program underperformed last month’s baseline because seasonality, creative mix, or offer quality moved.

Holdout vs A/B vs multi-variant

DesignQuestion it answers
A/B (A vs B)Which of these two is better relative to each other?
Multi-variant + predictionWhich page should this click see, given evidence so far?
HoldoutIs the new system better than the status quo we would have run anyway?

You can run a multi-variant race and keep a holdout on the old PDP or old campaign URL. Those answer different questions. See also A/B vs multi-variant.

Why holdouts matter for paid landing pages

Paid traffic is noisy: creatives rotate, auctions shift, promotions start and stop. Relative winners inside a closed race can look great while absolute ROAS falls.

A holdout answers: did introducing this landing system help versus doing nothing new?

Practical design notes

  1. Define the baseline — homepage, PDP, previous campaign page, or “ad → Shopify collection.”
  2. Size the holdout deliberately — large enough to estimate lift; small enough that the prediction system still gets learning traffic.
  3. Keep the holdout clean — do not quietly “improve” the baseline mid-test without logging it.
  4. Match attribution windows — compare purchases in the same window for both arms.
  5. Segment honestly — report by channel and major campaign; blended lift can hide Meta-only pain.
  6. Decide stop rules up front — calendar time, visit count, or confidence thresholds.

Common failure modes

How this shows up in Lutiq’s world

Conversion systems that predict per click still need a way to prove the program works. Holdout (or an equivalent baseline comparison) is how operators separate “one page looked good inside the experiment” from “we beat the destination we used before Lutiq.” Prefer pooled experiment vs holdout, not best-page vs holdout — see what A/B tests cannot see.

FAQ

Is a holdout the same as a control in A/B?

A control in A/B is usually one of the variants under test. A holdout is often the pre-existing production experience, kept aside from the optimization loop.

Should holdout traffic see message-matched pages?

Holdout should reflect the true alternative — typically the old stack. Matching the holdout to ads mid-test defeats the comparison.

How long should a holdout run?

Long enough to cover creative and promo cycles you care about. A three-day holdout during a flash sale rarely generalizes.


Part of Lutiq Learn. Analytics honesty before ranking theater.