learn-marketing-attribution-with-phoebe / Builder session 9 of 10
Learn Marketing Attribution with Phoebe · Builder session 9 of 10

Incrementality and geo-lift

Every model in this course so far - heuristics, Markov, Shapley, GA4 DDA, even MMM - is correlational. It tells you which channels showed up alongside the sale, never which channels caused it. This session builds the only causal leg of measurement: turn a channel off in some geos, keep it on in others, and measure the revenue delta. You design a geo-lift test for Lumen paid social, estimate the lift with a synthetic control, and then close the loop - feeding the measured lift back into the MMM as a Bayesian prior so the whole stack finally rests on something causal.

🔴 Hardest Builders experiments 45 min
0-4 · Why causal 4-20 · Holdout + synthetic control 20-40 · Test, estimate, calibrate 40-45 · Wrap
Part 0

Why causal is a different question

"Which channels were present when people bought?" and "which channels caused people to buy?" are different questions, and eight sessions of models have only answered the first. Correlation is enough for daily tactics; it is not enough to bet a budget. Incrementality is the only causal leg - and it is causal for one boring, powerful reason: it runs an experiment. It withholds a channel from a comparable group and measures what changes. Everything else in this course is a story fitted to observed data; this is the ground truth those stories get checked against.

Live - built in session Self-study - read after class ★ Build-along - run the code yourself Data: Lumen touchpoints + geo
★ What you walk out with today A powered geo-lift test design for Lumen paid social, a synthetic-control estimate of the lift with a real counterfactual, and the code that turns that lift into a Bayesian prior for the MMM - the move that upgrades your whole measurement stack from correlational to calibrated.
The one line that ties the course together MMM is correlational and needs incrementality experiments to be trusted. Lift-test results feed in as priors to recalibrate MMM elasticities. That sentence is this course's spine - today you build it.
Part 1 · the experiment

The holdout idea: on here, off there 6 min live

The whole design fits on a napkin. Split your markets into two comparable groups. Keep the channel running in the test geos; turn it fully off in the control geos. Watch both groups over the test window. The delta between what the test group did and what the control group did is the incremental effect - the sales that would not have happened without the channel.

TEST geos · channel ON US-CA · US-NY · US-TX paid social running full budget observed revenue = $412k CONTROL geos · channel OFF US-WA · US-IL · US-GA (comparable) observed revenue = $360k (scaled) Lift = delta between the groups $412k − $360k = $52k the incremental revenue paid social actually caused - not sales that would have happened anyway
🔍 Click to zoom - test geos keep the channel on, control geos turn it off, the revenue delta is the causal lift
LiveTwo flavours: geo holdout and platform RCT3 min

Same causal logic, two places to run it:

  • Geo holdout (geo-lift). You control the split - pick test and control markets, turn a channel off in the control geos, measure the delta. Works across any channel because you own the on/off switch. Meta GeoLift (R) is the reference tool, and it leans on synthetic control (Part 2).
  • Platform conversion-lift RCT. The ad platform randomly holds out a fraction of users from seeing your ads (Meta Conversion Lift), then compares the holdout to the exposed group. Cleaner randomization, but scoped to one platform and you trust their measurement.
Real world

A brand ran a Meta Conversion Lift study and found its retargeting was 80% incremental-fraud - the "conversions" were people who would have bought anyway. Last-touch had credited retargeting with a 6x ROAS; the holdout said the true incremental ROAS was near 1.2x. No correlational model could have caught that. Only turning it off could.

Self-studyQuasi-experiments when you cannot randomize2 min read

True randomization is not always possible - you cannot always cleanly turn a channel off, and geos are not identical. When randomization fails you fall back to quasi-experiments: designs that approximate a controlled test using observational structure. Synthetic control (Part 2) is the workhorse here - it builds a "fake" control group from a weighted blend of untreated geos when no single geo is a clean match. Difference-in-differences is its simpler cousin. These are weaker than a true RCT, but far stronger than any correlational attribution model.

Part 2 · the counterfactual

Synthetic control: building the counterfactual 6 min live

The hard part of any holdout is the counterfactual - what would the test geos have done if you had not run the channel? No single control geo matches perfectly. Synthetic control solves this: it builds a weighted blend of untreated "donor" geos that tracks the test group closely before the test starts, then projects that blend forward as the counterfactual. The gap between what actually happened and the synthetic line is the lift.

Actual test revenue vs synthetic control - the gap is the lift revenue test starts pre-period (fit the blend) test window (measure gap) lift actual test geos synthetic control (counterfactual) Lines must track closely BEFORE the test - that pre-period fit is what makes the post-period gap credible.
🔍 Click to zoom - synthetic control tracks the test group pre-test, then the post-test gap between actual and synthetic is the lift
LiveHow the weighted blend is built3 min

The counterfactual is a convex combination of donor geos - weights that are non-negative and sum to one - chosen so the blend matches the test group's pre-test trajectory as closely as possible:

  • Donor pool: all the geos that never got the treatment. These are the raw material for the synthetic control.
  • Fit on the pre-period: solve for weights that minimize the gap between the test group and the weighted donors before the test. A tight pre-period fit is the whole credibility of the method.
  • Project forward: hold those weights fixed and carry the blend through the test window. That projected line is what the test geos "would have done" - the counterfactual.
The credibility test Never trust a synthetic-control lift without seeing the pre-period fit. If actual and synthetic diverge before the treatment, the blend is wrong and the post-period gap is noise, not lift.
Self-studySpillover, and why the window is long2 min read

Two threats to a clean geo test:

  • Spillover: if your control geos border test geos, or your audience travels, the "off" group is not really off. Pick geographically separated markets and check for contamination.
  • Power and duration: a two-day test cannot detect a modest lift through the noise of daily revenue. You generally need 14 to 30-plus days so the accumulated signal clears the variance - and because adstock means the channel's effect takes weeks to fully express (Session 8).
Part 3 · the payoff

Closing the loop: lift tests calibrate MMM priors 5 min live

Here is why this session is not a detour. Your MMM from Session 8 estimated paid social's coefficient from correlation - it is a guess with a wide credible interval. A geo-lift test measures paid social's causal effect directly. In a Bayesian MMM you inject that measured lift as a prior on the coefficient, and the model tightens around the truth. The experiment does not replace the MMM - it calibrates it.

LiveWhy the loop beats either tool alone3 min
  • MMM alone covers every channel and every week but is correlational - it can be fooled by a channel that merely rode along with demand.
  • Lift tests alone are causal but expensive, slow, and scoped - you cannot run one on every channel every week.
  • The loop: run a few well-chosen lift tests, feed each as a prior into the Bayesian MMM, and the causal anchors pull the correlational estimates into line. The MMM then extends that calibrated truth across channels you never tested.
Real world

This is exactly how the big MMM vendors now work: Meridian and PyMC-Marketing both document experiment calibration as a first-class feature. A brand runs two or three geo-lift tests a year, uses them as priors, and the MMM's channel ROIs stop drifting on noise. The tests are the ground stakes; the MMM is the tent stretched between them.

Build-along 1 of 3

Design a geo-lift test for Lumen paid social ★ 8 min · run it

Before you touch a model, design the experiment: pick comparable geos, decide the split, and power the test so it can actually detect the lift you care about. A little Python turns "how long should we run this?" into a number.

Python · test designimport numpy as np import pandas as pd # Lumen revenue by geo x week (from the touchpoints + conversions tables) geo_wk = pd.read_csv("lumen_geo_weekly.csv", parse_dates=["week"]) # 1. candidate geos, ranked by how stable their weekly revenue is stats = (geo_wk.groupby("geo")["revenue"] .agg(mean="mean", std="std")) stats["cv"] = stats["std"] / stats["mean"] # lower = more predictable donors = stats.sort_values("cv").index.tolist() # 2. split: hold out a few comparable markets as CONTROL (channel off) test_geos = ["US-CA", "US-NY", "US-TX"] control_geos = ["US-WA", "US-IL", "US-GA"] # 3. power the test: weeks needed to detect a minimum lift we care about def weeks_for_power(baseline, weekly_std, mde=0.05, z=2.8): """mde = minimum detectable lift (5%); z ~ 80% power at 95% conf.""" effect = baseline * mde return int(np.ceil((z * weekly_std / effect) ** 2)) base = geo_wk.loc[geo_wk.geo.isin(test_geos), "revenue"].mean() sd = geo_wk.loc[geo_wk.geo.isin(test_geos), "revenue"].std() print(f"run for {weeks_for_power(base, sd)} weeks to detect a 5% lift")

Pick stable markets. Rank geos by coefficient of variation - predictable revenue makes a cleaner counterfactual. Wild, spiky markets drown the signal.

Comparable, separated split. Test and control groups should look alike historically but sit far enough apart that turning the channel off in control does not bleed into test (no spillover).

Power it. The weeks_for_power read tells you the honest test length for the smallest lift worth acting on. If it says 6 weeks, a 1-week test is theatre - it cannot see the effect.

Design before you spend A geo test costs real revenue in the control markets you go dark in. Powering it first means you never run a test too short to conclude anything - the most common and most expensive incrementality mistake.
Build-along 2 of 3

Estimate lift with a synthetic control ★ 8 min · run it

Build a weighted control from the donor geos that tracks the test group before the test, project it through the test window, and measure the gap. This is the core of Meta GeoLift, hand-rolled.

Python · synthetic controlfrom scipy.optimize import nnls # pivot to a weeks x geos revenue matrix mat = geo_wk.pivot_table(index="week", columns="geo", values="revenue", aggfunc="sum").sort_index() test_series = mat[test_geos].sum(axis=1) # actual test-group revenue donor_cols = [g for g in mat.columns if g not in test_geos] # split pre / post on the test start week start = "2026-05-04" pre = mat.index < start post = mat.index >= start # 1. fit non-negative weights on the PRE-period only (convex-ish blend) w, _ = nnls(mat.loc[pre, donor_cols].values, test_series[pre].values) w = w / w.sum() # normalise to sum to 1 # 2. project the synthetic control across the whole timeline synthetic = mat[donor_cols].values @ w synthetic = pd.Series(synthetic, index=mat.index) # 3. lift = actual minus synthetic, summed over the test window lift = (test_series[post] - synthetic[post]).sum() pre_fit_err = np.abs(test_series[pre] - synthetic[pre]).mean() print(f"pre-period avg gap (should be small): ${pre_fit_err:,.0f}") print(f"measured lift over test window: ${lift:,.0f}")

Fit on the past. nnls finds non-negative donor weights that reproduce the test group's pre-test revenue. Normalising to sum to 1 makes it a proper weighted blend.

Check the pre-fit first. Print pre_fit_err before you believe the lift. A small pre-period gap means the counterfactual is credible; a big one means stop - the blend does not match and the lift is noise.

Read the lift. The summed post-period gap is the incremental revenue paid social caused in the test geos - the causal number no correlational model could give you.

Real world

Meta's GeoLift package does exactly this with more rigour - it searches for the best test/control assignment, adds placebo tests (re-run the method on untreated geos to check the "lift" is really zero there), and gives a confidence interval. Your hand-rolled version is the honest core of it; the production tool is this plus guardrails.

Build-along 3 of 3

Feed the measured lift back as an MMM prior ★ 7 min · run it

Close the loop. Convert the measured lift into an expected coefficient for paid social, then set it as a prior in a Bayesian MMM so the model is anchored to a causal result instead of pure correlation.

Python · calibrate the prior# 1. turn the geo-lift result into an implied causal coefficient # lift revenue per unit of (adstocked, saturated) paid-social spend test_spend = geo_wk.loc[geo_wk.geo.isin(test_geos) & (geo_wk.week >= start), "paid_social_spend"].sum() implied_coef = lift / test_spend print(f"geo-lift implies paid_social coef ~ {implied_coef:.3f}") # 2. set it as an informative prior in a Bayesian MMM (PyMC-Marketing style) import pymc as pm with pm.Model() as mmm: # BEFORE: a vague prior - the MMM guesses from correlation alone # AFTER : centre the prior on the experiment, with modest uncertainty beta_social = pm.TruncatedNormal("beta_paid_social", mu=implied_coef, sigma=0.05, lower=0) # ... other channel coefs, adstock, saturation, likelihood ... # the causal prior pulls the fitted coefficient toward the tested truth print("MMM now calibrated: correlation tempered by a causal experiment")

Lift to coefficient. Divide the measured incremental revenue by the spend that produced it to get an implied causal coefficient for paid social - the number the experiment actually measured.

Prior, not fact. You do not hard-code it. You centre a prior on it with a sensible sigma, so the MMM blends the experiment with the rest of the data rather than ignoring either.

The loop is closed. The MMM's paid-social estimate is now anchored to a causal test. Run a few of these across your biggest channels and the whole model stops drifting on correlation. That is the calibrated stack Session 10 assembles.

Why this only works Bayesian A ridge or OLS MMM has nowhere to put the experiment - you can only fit the data. A Bayesian MMM has a prior slot built for exactly this outside knowledge. That is the real reason the field went Bayesian (Session 8): so the causal loop has somewhere to plug in.
Before Session 10

This week ◐ 45 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · Among MTA, Markov, Shapley, MMM and incrementality, which is the only causal method?

MTA, Markov, Shapley and MMM are all correlational - they describe what showed up alongside the sale. Only incrementality withholds a channel and measures what actually changes, which is why it is the sole causal leg.

2 · What does a synthetic control give you in a geo-lift test?

Synthetic control builds the counterfactual - a weighted combination of donor geos fitted to match the test group before the test, then projected forward. The gap between actual and synthetic is the lift. The pre-period fit is what makes it credible.

3 · How does an incrementality test improve an MMM?

The loop: a lift test measures a channel's causal effect, and that result becomes a prior on the MMM's coefficient. The correlational model gets anchored to a causal truth - which is why the field went Bayesian, so there is a slot for the experiment.

Source material

What this session covers

This session distills incrementality practice - holdout design, synthetic control, and MMM calibration - into a build-it-yourself pass on Lumen's geo data. Vendor tooling and full statistical treatments stay with their official sources.

Meta Blueprint - Conversion Lift pathholdout / RCT logic, incremental ROAS - Part 1
Meta GeoLift (R)synthetic control, geo test design - Parts 1-2, Build-alongs 1-2
Google Meridian / PyMC-Marketing - experiment calibrationlift as MMM prior - Part 3, Build-along 3
Synthetic-control theory (Abadie)named here; placebo tests in homework
The full unified stackassembled in Builder Session 10

Builder Session 9 cheat sheet · pin this

Incrementality =The only causal leg. MTA, Markov, Shapley, MMM are all correlational. Run an experiment: turn the channel off, measure the delta.
Holdout designChannel ON in test geos, OFF in comparable control geos. Lift = revenue delta - sales that would not have happened otherwise.
Synthetic controlWeighted blend of untreated donor geos, fitted to the test group's pre-period, projected forward as the counterfactual. Gap = lift.
Pre-fit is credibilityActual and synthetic must track BEFORE the test. Run a placebo on an untreated geo - the "lift" there should be ~0.
Power the test14-30+ days to clear the noise. Watch spillover between neighbouring geos. Design before you go dark - it costs real revenue.
Close the loopFeed measured lift into the MMM as a Bayesian prior. Causal test calibrates correlational model. This is the course's spine.