Glossary · revised 2026-07-13· defined with the product’s math
Randomized holdout
A randomized holdout is a group of leads selected by chance and deliberately never mailed. Because randomness — not judgment — decides who is withheld, the holdout’s outcomes reveal what would have happened without the program, and the gap between the mailed and withheld groups measures what the program actually caused. It is the design that separates “we sent email and deals arrived” from “the email caused the deals.”
The counterfactual, made observable
Every outbound program faces the same accounting problem: some of the leads it mails would have replied, booked, or bought anyway. Pipeline that appears after a campaign is a mixture of outcomes the campaign caused and outcomes that were already coming, and nothing observed inside the mailed group alone can separate the two.
Random withholding separates them. When chance assigns each lead to be mailed or withheld, the two groups are interchangeable in expectation before treatment — same industries, same buying intent, same seasonality. The withheld group then lives out the counterfactual: its reply and meeting rates are a direct observation of what would have happened anyway. Subtracting that baseline from the mailed group’s rate gives the program’s causal lift. Only random assignment makes the subtraction honest; a holdout chosen by judgment — say, the accounts a team was least excited about — builds that judgment into the baseline and biases the answer.
Why an A/B test cannot answer the same question
In an A/B test every arm is treated: subject line A goes to one half of the list, subject line B to the other, and no one is left unmailed. That design can rank messages against each other, but it contains no untreated group — no observation of the world without the program — so it cannot show that the program causes anything. A campaign whose arm A beats arm B by a wide margin can still be producing zero incremental pipeline if the replies were coming regardless.
The distinction is settled practice in adjacent channels. Lifecycle-email platforms ship holdouts natively — Customer.io’s workflow documentation (updated 10 July 2026) describes a holdout branch that deliberately withholds the message while conversions keep being tracked for the people who never received it — and incrementality vendors such as Measured formalize the same design for email, catalog, and SMS as a known-audience split. Mainstream cold-email platforms show no native holdout in their published docs and changelogs as of July 2026; the channel’s tooling reports deliverability and reply rates on the mailed population only.
What a holdout costs, and how it is sized
A holdout is not free: every withheld lead is a prospect the program deliberately does not work. Sizing one is an underwriting decision — a larger withheld share buys evidence faster at the price of more forgone outreach, and the required size depends on the baseline rate and on the smallest effect worth detecting.
RevenueOS prices this openly. Its minimum-n table (generated 3 July 2026 from the shipped proof engine, not from a textbook power formula) puts the floor for one mid-range scenario — a 0.1% holdout baseline against 1.2% positive replies per mailed lead — at a median 9,260 evidence rows before the 90% anytime-valid confidence sequence (α = 0.1) narrows to the planned effect size τ.
How RevenueOS uses this
RevenueOS randomizes leads, never emails: each lead is assigned once, at enrollment, to a cohort, so a withheld lead stays withheld across every step of every sequence and cannot leak into the mailed population mid-experiment. The shipped design withholds 10% of enrolled leads and mails a matched 10% proof cell; the published contrast reads only those two cells. Inside the proof cell the message arms are rotated uniformly — deliberately not Thompson-optimized — so the published claim is exact about its estimand: the arm menu, uniformly rotated, causes X versus not mailing. That reads as a lower bound for the optimized program, under the stated assumption that the optimizer does no worse than uniform rotation.
Partners who need evidence sooner can widen the split — the same table shows a 25%/25% design accruing evidence rows 3.0× faster than 10%/10% at the same sends budget. RevenueOS is at design-partner stage as of July 2026; the numbers above are design floors and simulation output, not customer results.
Related terms
Incrementality · Causal lift · Thompson sampling · Minimum detectable effect
Sources
- Customer.io Docs — Holdout tests in campaign workflows (page updated 10 Jul 2026)
- Measured — What is incrementality in marketing? (Known-Audience Split for email, catalog, and SMS)
the full statistical machinery, in writing: /methods · every term: /glossary