Guide · revised 2026-07-13 · vendor claims cited from their own docs

Attribution vs. incrementality: credit is not proof

Attribution assigns credit backward: it splits revenue among the touches a buyer was seen to have, in proportions a model chooses. Incrementality proves cause forward: it withholds a randomized group and measures what happens without the touch. The two answer different questions — which touches co-occurred with the deals you won, versus how many of those deals would have happened anyway — and only the second can justify a budget.

What attribution actually computes

Every attribution model — first-touch, last-touch, U-shaped, W-shaped, data-driven — starts from the same raw material: the journeys of buyers who converted, and the touches observed along the way. The model then distributes credit across those touches by rule or by fitted weights. Nothing in that procedure ever observes a world where a touch didn’t happen, so nothing in it can distinguish a touch that caused the deal from a touch that merely attended it.

The candid vendors say this themselves. Dreamdata’s own documentation explains that its data-driven model looks for correlations, and offers its own example: a meeting logged before every closed deal will earn heavy credit, yet the meeting is probably not causing the sales — and the problem worsens in B2B’s long, non-linear, multi-stakeholder journeys. HockeyStack’s docs put it in four words: multi-touch attribution is “purely directional, not incremental.”

The failure mode for outbound is concrete: a hot inbound account that a rep also cold-emailed converts, and the attribution model pays the cold email a share of the credit — for a deal that was closing anyway. Credit is a bookkeeping convention. It is not evidence.

What incrementality actually measures

An incrementality measurement builds the missing world deliberately: before anything is sent, leads are randomized into a treated group and a randomized holdout that never receives the program. Because assignment was random, the holdout’s outcomes estimate what the treated group would have done untouched, and the difference — the causal lift — is attributable to the program itself, selection bias removed by construction.

The word “randomized” is load-bearing. Comparing touched accounts against accounts that happened to go untouched — the retroactive “lift” some analytics products compute — quietly re-imports the bias, because reps touch the accounts they judged most likely to convert. Randomization is the only thing that removes that judgment from the comparison.

Where the peeking problem comes in

A holdout alone is not enough, because outbound results arrive as a trickle — a reply today, a meeting next week — and everyone reads the numbers as they land. Classical fixed-sample statistics are only valid at one pre-committed look; checked daily, a 95% confidence procedure’s real false-positive rate climbs far above its label. That is the peeking problem, and it silently voids most informal holdout readings.

The repair is to use statistics built for continuous reading: anytime-valid confidence sequences hold their guarantee at every look simultaneously, so watching the experiment daily is legitimate rather than corrosive. The full machinery, sample-size floors included, is written down at /methods.

A worked example

The numbers here are illustrative planning arithmetic, not a claim about any live program. A program mails 9,000 leads in a quarter and holds out 1,000 at random. The treated group books 36 meetings — 0.40%. The holdout, never mailed, books 1 through other channels — 0.10%. Attribution credits the sequence with all 36, since the touch appears on every journey. The holdout says otherwise: at a 0.10% baseline, about 9 of those 9,000 treated leads were coming anyway, so the program’s causal contribution is roughly 27 meetings — three-quarters of the credited figure. And at these sample sizes the confidence band around that difference is still wide, which is exactly why published floors and anytime-valid bands exist instead of a point estimate and a victory lap.

When an attribution dashboard is the right tool

Honesty cuts both ways: attribution is the right instrument for real jobs. Mapping how buyers actually move, triaging budget across a dozen simultaneous channels, spotting a channel whose trend broke — those are journey-description problems, and a well-built attribution platform describes journeys better than any experiment will. It is also the honest fallback when you cannot withhold: a five-account target market where every named account matters, or a decision too small to be worth an experiment’s cost.

Even the vendors draw the line roughly where we do. Dreamdata’s incrementality page frames formal testing as heavy machinery for the largest budgets, and HockeyStack positions its lift product as the validator layered over its own attribution. The working rule: use the dashboard to generate hypotheses; use a randomized holdout before a hypothesis is allowed to spend real money.

the register: if a number decides budget, it should come from a design that could have proven you wrong. · the tools, ranked: cold email testing tools in 2026 · the terms, defined: /glossary