learn-experimentation-with-phoebe / Leader session 4 of 6
Learn Experimentation with Phoebe · Leader session 4 of 6

Beyond the button, when you can't A/B a person

A clean user-level A/B test is a luxury. A TV spot, a billboard, an out-of-home burst, a privacy-limited channel, a brand campaign that spills across everyone - none of these can be split shopper by shopper. Yet these are often your biggest bets. This session gives you the leader's toolkit for causal answers when the button test is off the table: holdouts and incrementality, geo experiments, and why last-click reporting quietly overstates your paid channels. No math, no code - just knowing which tool fits which question.

🟡 Core Leaders · C-level / PM / marketing No code 45 min
0-3 · Setup 3-22 · Holdouts & incrementality 22-40 · Geo experiments 40-45 · Wrap
Part 0

The channels you cannot split by shopper

Everything you learned in Sessions 1 to 3 assumed you could flip a coin per person: this shopper sees the new page, that one sees the old. But some of Lumen's most expensive moves cannot be assigned that way. A connected-TV burst reaches whole households. A billboard is seen by a city, not a cookie. Privacy limits mean you can no longer follow individuals across apps and browsers. And a brand campaign lifts everyone at once, so there is no untouched control person left. When the unit of a test cannot be a single user, leaders still need a defensible answer to the only question that matters: would this have happened anyway?

Live - covered in session Self-study - read after ★ The one big idea The case: Lumen Skincare
★ What you walk out with today Three tools and the judgement to pick between them. A holdout deliberately withholds a channel from a slice of the audience so you can measure its true incremental lift. A geo experiment splits by region instead of by person - some markets get the campaign, the rest are the holdout. And incrementality is the one honestly causal number in marketing: the sales that would not have happened without the spend. You will leave able to tell when last-click reporting is flattering a channel, and when to ask for a geo test instead of a dashboard.
Part 1 · the one big idea

Incrementality is the only causal number in marketing 8 min live

Lumen's dashboard says paid search drove $1.2M in revenue. It says so because paid search was the last click before the sale. But most of those shoppers already wanted Lumen - they typed the brand name, clicked the ad sitting on top of the free result, and bought. Would they have bought anyway? Last-click reporting cannot answer that; it hands full credit to whatever touch happened last. The honest question is incremental: how many of those sales would not have happened without the spend. The only way to find out is to withhold the channel from a comparable group and watch the gap - a holdout.

Reported vs incremental: paid search on Lumen Reported (last-click) $1.2M Credited to the last touch before the sale True incremental ~$0.4M Sales that would NOT have happened without it The gap = budget on sales you'd get anyway Last-click reporting reliably overstates paid channels shoppers were going to use anyway. Incrementality is what a holdout actually measures.
🔍 Click to zoom - reported credit is not incremental lift; the gap is wasted budget
LiveWhat a holdout actually does3 min

A holdout is a control group for a whole channel. You deliberately withhold the ads from a randomly chosen slice of a comparable audience, run everything else the same, and compare. The difference in sales between the exposed group and the held-out group is the channel's true incremental contribution.

  • Reported revenue - what the platform or last-click dashboard credits to the channel. Flattering, and usually too high.
  • Incremental revenue - the sales that would not have happened without the spend. The only number worth reallocating budget on.
  • The holdout - the comparable group you withheld the channel from. Your window into "what would have happened anyway", exactly like Session 1's counterfactual.
Real world

Lumen holds paid search back from a random slice of its audience for four weeks. Sales in the held-out slice barely move. The lesson lands hard: much of that $1.2M was demand Lumen already had. This is the same trap as Session 1's email story - a channel that looks like a hero because high-intent shoppers pass through it, not because it created the intent.

Self-studyThe bridge to attribution: MMM needs an experiment to trust it4 min read

If you have taken the sibling learn-marketing-attribution course, this session is where the two meet. Attribution models - last-touch, data-driven, and marketing-mix modelling (MMM) - are all correlational. They divide up credit; they do not prove cause. Only incrementality does.

  • MMM (that course's B8) reads years of spend and revenue to estimate each channel's contribution. Powerful, but it can drift - it is fitting curves to observational history.
  • Geo-lift and incrementality (that course's B9) supply the ground truth. A geo experiment gives one honest causal number that the MMM is then calibrated against - the experiment anchors the model so its channel estimates stay believable.
  • The leader takeaway: when a vendor shows you an MMM, ask "what experiment is this calibrated to?" A model with no experiment behind it is a confident guess.
One line to remember Attribution tells you where credit went. Incrementality tells you what your money actually caused. Never reallocate a $4M budget on attribution alone.
Part 2 · the leader's tool

Geo experiments: split by region, not by person 7 min live

When you cannot flip a coin per shopper, flip it per region. A geo experiment turns whole markets into the units of the test: one set of regions gets the campaign, the rest are held out as the comparison. It is the natural home for exactly the channels a button test cannot reach - CTV, out-of-home, brand bursts - because those reach a place, not a cookie. Lumen's data already carries 9 US regions, which makes it a ready-made geo lab.

User-level split (Sessions 1-3) One coin flip per shopper. Impossible for a TV spot or billboard - it reaches everyone. Geo-level split (this session) W-1 C-1 E-1 W-2 C-2 E-2 W-3 C-3 E-3 Treated: gets the CTV burst Held out: the comparison The 9 held-out and treated regions must look alike before the campaign. The lift in treated markets over held-out ones is the incremental effect.
🔍 Click to zoom - regions become the units; treated markets minus held-out markets is the lift
LiveWhat a geo test can and cannot tell you3 min

A geo experiment is powerful, but it is a blunter instrument than a user-level A/B. Knowing its limits is the leader's job.

  • Can: give a credible causal read on channels that reach a place, not a person - CTV, radio, out-of-home, a brand campaign. It answers "did the burst lift sales in these markets versus the rest?"
  • Can: anchor an MMM. One good geo test calibrates the whole model.
  • Cannot: give you many units. You have 9 regions, not 32,000 shoppers, so a geo test is lower-powered and needs a bigger, longer effect to detect. Small tweaks will not show up.
  • Cannot: escape the assumption that treated and held-out regions were tracking together before you started. If the West was already booming, the "lift" is contaminated - which is exactly the parallel-trends worry Session 5 unpacks.
Real world

Lumen commissions a geo test for a connected-TV burst: 4 of its 9 US regions get the CTV campaign, 5 are held out. Over 8 weeks, treated markets run about 3.6% conversion against 3.2% in the held-out markets - the same 0.4 point gap as the flagship button test, but now proven for a channel no cookie could track. That single number is what the CMO uses to defend the CTV line in the $4M budget.

Self-studyWhich tool for which question3 min read

A leader does not run these - you commission the right one. The whole skill is matching the question to the tool:

The situationThe right tool
A page or flow change you can serve per shopperUser-level A/B test (Sessions 1-3)
"Is this paid channel actually incremental?"Holdout test
TV, CTV, radio, out-of-home, a brand burstGeo experiment
"Where should the whole $4M go?"MMM, calibrated by a geo test
A change already rolled out to some regions, no clean controlQuasi-experiment (Session 5)
Your leverage is the commission, not the calculation When someone brings you a brand-campaign result with no holdout and no geo split, the right question is not "what is the ROAS?" It is "what was the comparison group, and would these sales have happened anyway?"
Before Session 5

This week ◐ 30 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · Your dashboard credits paid search with $1.2M in revenue. Why is that not the amount to reallocate budget on?

Last-click reporting hands full credit to the final touch. Only a holdout measures incrementality - the sales that would not have happened without the spend, which is usually far lower.

2 · Lumen wants a causal read on a connected-TV burst. Why is a user-level A/B test the wrong tool?

When the treatment reaches a place rather than a cookie, the unit of the test has to be the place. A geo experiment treats some of the 9 regions and holds the rest out.

3 · A vendor presents an MMM saying influencer spend should double. What is the sharpest leader question?

MMM is correlational - it fits curves to historical spend and revenue. A geo experiment supplies the one causal number that anchors it. An MMM with no experiment behind it is a confident guess.

Source material

What this session covers

This leader session distills the causal-marketing toolkit that the builder track and the sibling attribution course build in code. It covers ~80% of the conceptual content; the mechanics live in the sources below.

Holdouts and incrementality - the only causal number in marketingexec framing - Part 1
Last-click overstates paid channels; reported vs incrementalthe gap diagram - Part 1
Geo experiments - split by region, what they can and cannot tell youthe leader's tool - Part 2
Meta Marketing Analytics / GeoLift, Google Geo experiments (Vaver-Koehler)the platforms named, not operated
Geo holdout, donor pool, CausalImpact in PythonBuilder Session 8, and learn-marketing-attribution B9
MMM calibration (Meridian / PyMC-Marketing)learn-marketing-attribution B8 - anchored by this session's geo test

Leader Session 4 cheat sheet · pin this

IncrementalityThe sales that would NOT have happened without the spend. The only causal number in marketing. Reallocate on this, not on reported credit.
Holdout testWithhold a channel from a random comparable slice, compare sales. The channel-level version of a control group.
Last-click overstatesReporting hands full credit to the final touch, so channels shoppers use anyway look like heroes. The gap to incremental is wasted budget.
Geo experimentSplit by region, not by person. Treat some markets, hold the rest out. The home for CTV, TV, out-of-home, brand bursts.
Geo limitsFew units (9 regions), so lower power - only sizable effects show. Assumes treated and held-out markets tracked together before.
MMM needs an experimentAttribution and MMM are correlational. A geo test supplies the causal anchor. Ask "what experiment is this calibrated to?"