The working vocabulary of causal proof for cold outbound. Each entry answers in its first breath, shows the mechanics with the product’s actual math, and cites dated primary sources — because a definition page that carries no substance is not worth a machine’s citation or a reader’s minute. The full statistical machinery lives on the methods page.
A randomized holdout is a group of leads selected by chance and deliberately never mailed. Because randomness — not judgment — decides who is withheld, the holdout’s outcomes reveal what would have happened without the program, and the gap between the mailed and withheld groups measures what the program actually caused. It is the design that separates “we sent email and deals arrived” from “the email caused the deals.”
Incrementality is the portion of outcomes that would not have happened without the action — the replies, meetings, or revenue a campaign caused, over and above what the same audience would have produced anyway. It is a forward-looking causal quantity: it can only be measured by deliberately withholding the action from a randomized group and comparing, not by assigning credit after the fact among touches that all occurred.
Causal lift is the change in an outcome rate caused by a treatment — the treatment effect, written τ: the rate a cohort achieved when treated minus the rate the same cohort would have achieved untreated. Only a randomized contrast identifies it, because randomization is what makes the untreated group a fair stand-in for the counterfactual. Lift computed retroactively from exposed-versus-unexposed groups measures selection as much as effect.
An anytime-valid confidence sequence is a confidence interval you may legally read at any time: its coverage guarantee holds uniformly over every sample size and every stopping rule, so continuous monitoring costs nothing statistically. A fixed-n interval promises coverage only at one preplanned look — the moment you peek and act on what you see, that guarantee is void. A sequence run at α = 0.1 keeps 90% coverage however often anyone checks.
Peek at a running A/B test whenever the dashboard refreshes, stop the moment it shows significance, and the false-positive rate climbs far above the nominal α the test promised. That is the peeking problem: a fixed-horizon significance test is valid only at one preplanned look, and optional stopping — deciding when to stop based on the data so far — voids its guarantee.
Thompson sampling is an allocation rule for sequential experiments: assign the next trial to each competing option with probability equal to the current posterior probability that it is the best one. In practice, draw one plausible value per arm from its posterior distribution and play the arm with the largest draw. William R. Thompson proposed the rule in Biometrika in 1933, decades before the multi-armed-bandit literature adopted it.
The minimum detectable effect (MDE) is the smallest true effect a given experimental design can reliably distinguish from zero at its planned sample size and confidence level. It is the budget line of experimentation: an effect below the MDE can be real and still produce a null read, so a null result from an underpowered design is uninformative, not exonerating. Honest testing states its MDE before the first send.
Positive reply rate is the share of leads contacted who send a genuinely interested reply — a question, a referral, or an agreement to talk. It excludes unsubscribes and plain refusals, and its denominator is leads, not emails sent and not replies received. Confusing it with raw reply rate, or with the share of replies that are positive, moves the number by roughly an order of magnitude.