learn-experimentation-with-phoebe / Leader session 5 of 6
Learn Experimentation with Phoebe · Leader session 5 of 6

When you can't randomize, and still need an answer

Sometimes there was no experiment. The loyalty program already launched in three regions. The policy already changed. Nobody flipped a coin, and you cannot rewind. You are left with observational evidence - and the temptation to read it as if it were a clean test. This session gives leaders the two most trusted quasi-experimental tools, difference-in-differences and synthetic control, in plain language, plus the one caveat that keeps you honest: these methods only correct for the things you thought to measure. No math, no code.

🟠 Advanced Leaders · C-level / PM / marketing No code 45 min
0-3 · Setup 3-22 · Difference-in-differences 22-40 · Synthetic control 40-45 · Wrap
Part 0

The world that already happened

A randomized test is the gold standard because the coin flip guarantees the two groups are alike in every way, measured or not. Quasi-experiments give that up. They work with the world as it unfolded - a program rolled out to some regions first, a change that hit one market and not another - and try to reconstruct the missing counterfactual after the fact. They can be genuinely convincing. They can also be quietly wrong. The leader's job is to know which, and to hold the evidence to the right standard.

Live - covered in session Self-study - read after ★ The one big idea The case: Lumen Skincare
★ What you walk out with today Two tools and one warning. Difference-in-differences compares the change in a treated group to the change in a comparison group, so shared trends cancel out. Synthetic control builds a bespoke "Lumen-without-the-campaign" from a weighted blend of untreated regions, and judges its own credibility by how well it matches the real Lumen before the change. The warning: both only adjust for confounders you measured - the unknown ones remain, which is why these are weaker than a randomized test and should carry real humility.
Part 1 · the one big idea

Difference-in-differences: compare the changes, not the levels 8 min live

Lumen rolls a new loyalty program out to some regions before others. You cannot compare the treated regions' sales to the untreated ones directly - they were never identical to begin with. But you can compare how much each group changed. If the untreated regions rose 5% over the period and the treated regions rose 12%, the extra 7 points is the loyalty program's effect - assuming the two groups would have moved together without it. That subtraction of two changes is the whole idea, and it is why it is called difference-in-differences.

weekly revenue → loyalty launch Control Treated would-have path gap = the effect Before launch the two rise in parallel - the assumption you must check. After launch, treated pulls above its would-have path.
🔍 Click to zoom - parallel before, a gap after; the gap over the would-have path is the effect
LiveThe parallel-trends assumption, in plain words3 min

Difference-in-differences rests on a single "what if": would the two groups have moved together if nothing had changed? That is the parallel-trends assumption, and it is the whole ballgame.

  • What it claims - absent the loyalty program, treated and comparison regions would have risen and fallen by the same amounts. Their lines would have stayed parallel.
  • How you check it - look at the period before the change. If the two groups tracked together for months beforehand, the assumption is credible. If they were already diverging, it is not.
  • How it breaks - if Lumen deliberately launched the program first in its fastest-growing regions, those were already pulling ahead. The "effect" would just be that head start.
Real world

Lumen launches loyalty in 4 of its 9 regions. In the year before launch, treated and untreated regions grew almost in lockstep - good news for the assumption. After launch the treated regions accelerate, opening a gap the analysts read as roughly a few points of extra conversion. Because the pre-period was clean, the CMO trusts it enough to fund a wider rollout - but commissions a proper geo experiment to confirm.

Self-studyWhy the double subtraction is clever3 min read

The elegance is that subtracting two changes cancels out anything shared. A nationwide promo, a seasonal dip, a supply shock - if it hit both groups equally, it lifts or drops both changes by the same amount and washes out of the difference. That is why difference-in-differences can survive a messy year that a simple before-and-after cannot.

  • Before-and-after (one group) - tangles the change with every other thing that moved. Session 1's trend trap.
  • Treated vs control levels (one moment) - tangles the effect with the fact that the groups were never identical.
  • Difference-in-differences - subtracts both problems away, as long as the shared-trends assumption holds.
The leader question "Show me the pre-period." A difference-in-differences result with no picture of the before-launch trends is asking you to take the key assumption on faith.
Part 2 · the leader's tool

Synthetic control: build a fake Lumen that never got the campaign 7 min live

Sometimes only one region got the treatment, and no single other region is a good match for it. Synthetic control solves this by building a made-up comparison out of a weighted blend of the untreated regions - say 40% of the East, 35% of the Central, 25% of the West - tuned so the blend traces the treated region's history almost exactly before the campaign. That blend becomes your "Lumen-without-the-campaign". After the campaign, the gap between the real region and its synthetic twin is the estimated effect.

weekly revenue → campaign start Real Synthetic gap = the effect Synthetic twin = 40% East + 35% Central + 25% West A tight pre-campaign match (the two lines sitting on top of each other) is the credibility test. A poor match means do not trust the gap.
🔍 Click to zoom - a good pre-period match earns the right to read the post-period gap
LiveThe pre-period match is the credibility test3 min

Synthetic control has a built-in honesty check that leaders can read without any math: how well does the synthetic twin match the real region before the campaign?

  • Tight match - the synthetic line sits almost on top of the real one for the whole pre-period. That earns the right to believe the post-period gap is real.
  • Loose match - if the blend could not even reproduce the past, it has no business predicting the counterfactual. Distrust the effect.
  • Placebo check - good analysts also run the method pretending an untreated region was treated. If those fake "effects" are as big as the real one, the real one is probably noise.
Real world

Lumen ran its CTV burst hardest in one standout region. No single other region matched it, so the analysts built a synthetic twin from a blend of the other 8. The twin tracked the real region within a whisker for the full year before the burst, then the real region pulled clearly above it during the campaign. Because the pre-period fit was that tight, the result held up in the budget review.

Self-studyThe honest caveat: only measured confounders3 min read

Here is the line that separates a careful leader from an overconfident one. Both difference-in-differences and synthetic control can only adjust for the things you measured. A randomized test balances even the confounders nobody thought of, because the coin flip does not care what is measured. Observational methods cannot - an unknown factor that hit only the treated group will masquerade as effect.

  • Randomized test - balances known and unknown confounders alike. Strongest evidence.
  • Quasi-experiment - balances only what you modelled and the assumption you made. Good, not gold.
  • The humility rule - present the effect with its assumption stated out loud, and treat it as strong evidence, not proof.
When to trust it vs demand a real experiment Trust a quasi-experiment when a randomized test is genuinely impossible, the pre-period is clean, and the stakes are moderate. Demand a real experiment when the decision is large and hard to reverse and a randomized or geo test is feasible - the humility of the method should scale with the size of the bet.
Before Session 6

This week ◐ 30 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · Lumen launched loyalty in some regions first. Why compare the change in each group rather than the raw sales levels?

Difference-in-differences subtracts the comparison group's change from the treated group's change. Anything shared - seasonality, a national promo, the initial level gap - washes out, leaving the treatment effect.

2 · What single thing would most convince you a synthetic-control result is credible?

If the blend cannot even reproduce the region's own history, it cannot be trusted to predict the counterfactual. A tight pre-period match is the credibility test; placebo checks reinforce it.

3 · Why is a quasi-experiment weaker than a randomized test, even with a clean pre-period?

Randomization balances known and unknown confounders alike because the coin flip does not care what was measured. Observational methods correct only for what you modelled, so they carry an assumption a randomized test does not.

Source material

What this session covers

This leader session distills the quasi-experimental toolkit the builder track builds in code. It covers ~80% of the conceptual content; the mechanics live in the sources below.

Difference-in-differences and the parallel-trends assumption, in plain languageexec framing - Part 1
Synthetic control - donor blend, pre-period match, placebo checksthe leader's tool - Part 2
The observational caveat - only measured confounders, when to trust vs demand a testthe humility rule - Part 2
DiD in code - TWFE, event study, staggered-adoption critiqueBuilder Session 9
Synthetic control (Abadie) and CausalImpactBuilder Session 8
UPenn Crash Course in Causality; Resende Econometrics (Udemy)the formal methods and code home

Leader Session 5 cheat sheet · pin this

Quasi-experimentReconstructing the counterfactual from the world as it happened, when no coin was flipped. Convincing but weaker than a randomized test.
Difference-in-differencesCompare the CHANGE in a treated group to the CHANGE in a comparison group. Shared trends cancel out.
Parallel trendsThe key assumption: the two groups would have moved together without the change. Check it in the pre-period. Ask "show me the before."
Synthetic controlBuild a fake "Lumen-without-the-campaign" from a weighted blend of untreated regions. Best when only one unit was treated.
Pre-period matchThe credibility test - a tight fit before the change earns the right to read the gap after it. Placebo checks confirm.
The humility ruleThese adjust only for confounders you measured. State the assumption out loud; demand a real experiment when the bet is big and reversible tests exist.