learn-experimentation-with-phoebe / Leader session 1 of 6
Learn Experimentation with Phoebe · Leader session 1 of 6

Why experiments beat opinions, and dashboards

You do not need to run a test to lead a team that runs great tests. You need to know what a real experiment buys you that a strong opinion and a rising chart do not: a defensible answer to "did this actually cause the result?" This session gives you the one idea underneath the whole discipline - the counterfactual - and a clear rule for when a decision is worth an experiment at all. No math, no code. This is your ▶ start button for the leader track.

🟢 Start here Leaders · C-level / PM / marketing No code 45 min
0-3 · Setup 3-22 · The counterfactual 22-40 · When to commission 40-45 · Wrap
Part 0

The most expensive sentence in business

"We shipped it, and the numbers went up, so it worked." That sentence has justified more bad budget decisions than any spreadsheet error. It sounds like evidence. It is a story - one that quietly assumes the numbers would have stayed flat if you had done nothing. This session is about the missing half of that sentence: the world where you did not ship it. Get comfortable with that missing world - the counterfactual - and you can tell a genuine result from a lucky one, and know which decisions deserve the cost of a real test.

Live - covered in session Self-study - read after ★ The one big idea The case: Lumen Skincare
★ What you walk out with today A leader's working model of causation: what "it caused the lift" actually claims, why correlation and before/after charts routinely fool smart people, and a one-page rule for deciding when a decision is big or reversible enough to be worth an experiment - and when it plainly is not. You will be able to sit in a results review and ask the two or three questions that separate a real win from a mirage.
Part 1 · the one big idea

Causation is a comparison to a world you cannot see 8 min live

When Lumen's team says the new product page "caused" a lift, they are making a claim about two worlds: the one where these shoppers saw the new page, and the one where the very same shoppers, same day, saw the old one. Only one of those worlds ever happens. The gap between them is the true effect - and the entire craft of experimentation is a set of honest tricks for estimating a world you never got to observe.

The tempting story: before → after Last month3.2% convert This month3.6% convert "+0.4pt, it worked!" - but a promo, payday, and the season all changed too. The trend is not the effect. The honest test: same time, split traffic Control (old)3.2% convert Variant (new)3.6% convert 🎲 Same shoppers, same day, split by coin flip. The only difference is the page - so the gap IS the effect. The control group is your window into the counterfactual: it is "these kinds of shoppers, this week, without the change." Before/after has no control group, so it cannot separate what you did from everything else that moved. A rising chart after a launch is not proof. Proof needs a comparable group that did not get the change.
🔍 Click to zoom - before/after tells a story; a control group estimates the counterfactual
LiveCorrelation, causation, and the confounder3 min

Ice-cream sales and drowning rise together. Neither causes the other - summer heat drives both. That hidden common cause is a confounder, and it is why "X and Y move together" almost never means "X causes Y".

  • Correlation - two things move together. Cheap, everywhere, often coincidental.
  • Causation - changing one changes the other. Expensive to prove, and the only thing worth betting budget on.
  • Confounder - a third factor driving both, manufacturing a correlation with no causal link.
Real world

Lumen finds its email subscribers convert three times better than non-subscribers and nearly doubles the email budget. But subscribers chose to subscribe - they were already high-intent fans. The 3x is mostly who they are, not what email did. Engagement is the confounder. A holdout test later showed email's true incremental lift was a fraction of that. This is the single most common way marketing budgets get misallocated.

LiveWhy smart, honest people still get fooled4 min

It is not stupidity - it is that the wrong methods feel exactly as convincing as the right ones. Three traps catch everyone:

  • The trend trap - attributing a seasonal or promo-driven rise to your change (the before/after picture above).
  • The self-selection trap - comparing people who opted in to people who did not (the email story).
  • The peeking trap - watching a live test daily and declaring victory the first day it looks good. The tool below shows how badly this one lies - and it feels like diligence, not error.

You do not need to run the statistics to defend against these. You need to recognize the shape of each and ask for the control group, the randomization, and the pre-committed stopping point.

See the peeking trap for yourself

Watching a test every day manufactures false wins try it

Here is a test where the two versions are genuinely identical - there is no real difference at all. Run it once checking only at the end, then run it checking every day and stopping the first time it looks significant. The "checking every day" version declares a false winner far more than the honest 5% of the time. Nobody lied; they just looked too often. This is why a good experimentation culture pre-commits to a stopping point.

Part 2 · the leader's judgement

When is a decision worth an experiment? 6 min live

Experiments cost traffic, time, and focus. Not every decision earns one. Your job as a leader is not to test everything - it is to spend your experimentation budget where the stakes and the uncertainty are both high, and to move fast on the rest. Here is the map.

high stakes / costly to undo → low uncertainty → high uncertainty Just ship it Low stakes, easy to reverse, you are fairly sure. Copy tweak, button colour. Testing costs more than being wrong. Move. Run a quick test Low stakes but you honestly do not know. A light A/B settles it cheaply and builds the habit. Bias toward testing here. Pilot / staged rollout High stakes, hard to undo, but you are confident. Roll out to one region first, watch guardrails, then expand. A geo test (Session 4) is the tool. Experiment - always High stakes AND high uncertainty. Pricing, checkout flow, a big brand bet. This is exactly what your experimentation budget exists for. Never guess here. Reversibility matters as much as size: a cheap-to-undo mistake is a fast lesson; an expensive-to-undo one needs evidence first.
🔍 Click to zoom - stakes x uncertainty decides whether to ship, pilot, or experiment
Self-studyThe three questions to ask in any results review3 min read

You will spend more time reading experiments than commissioning them. These three questions, asked out loud, catch most bad claims - and you will have the vocabulary for all three by the end of the leader track.

  • "What is the control group?" - if there is no comparable group that did not get the change, it is a before/after story, not a test. (Session 1, today.)
  • "Was it powered, and did we pre-commit to the sample size?" - underpowered tests and peeked-at tests both produce fake wins. (Session 2 and 3.)
  • "Did anything about the split look off - lopsided groups, a weird week?" - the trust checklist that separates a clean test from a broken one. (Session 3.)
Your leverage is the question, not the calculation A leader who reliably asks "where is the control group and did we pre-commit the sample size?" raises the quality of every experiment their org runs - without opening a notebook. That is the whole point of this track.
Before Session 2

This week ◐ 30 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · A conversion chart rises the week after a launch. Why is that not proof the launch caused it?

Before/after has no control group, so seasonality, promos, and other shifts are baked into the "lift". The control group is your estimate of the counterfactual - what would have happened anyway.

2 · Email subscribers convert 3x better than everyone else. The safest reading is...

Self-selection makes the groups non-comparable. Only a holdout/randomized test isolates email's true incremental effect, which is usually far smaller than the raw gap.

3 · Which decision most clearly deserves a full experiment?

Experimentation budget belongs where stakes AND uncertainty are both high. Low-stakes, easy-to-reverse, confident changes should just ship - testing them costs more than being wrong.

Source material

What this session covers

This leader session distills the "why" that the technical courses assume you already believe. It covers ~80% of the conceptual foundation; the mechanics live in the builder track and the sources below.

Causation vs correlation, the counterfactual, confoundingexec framing - Part 1
When to commission an experiment (stakes x uncertainty)the decision grid - Part 2
Kohavi, Tang & Xu - Trustworthy Online Controlled Experimentsthe industry bible; we draw the culture chapters in Session 6
Brady Neal - Intro to Causal Inference (concepts)the formal counterfactual model - Builder Session 1
The statistics (power, p-values, SRM)Leader Sessions 2-3, no math required

Leader Session 1 cheat sheet · pin this

The counterfactualCausation = comparing what happened to what would have happened without the change. You never see it directly.
Control groupYour window into the counterfactual: comparable people who did not get the change, same time.
ConfounderA hidden common cause that fakes a correlation. Self-selection and seasonality are the usual suspects.
Before/after ≠ proofA rising chart after a launch tangles your change with everything else that moved. Ask for the control group.
Peeking liesChecking a live test daily and stopping at first "win" manufactures false positives. Pre-commit the sample size.
When to testHigh stakes AND high uncertainty → experiment. Low stakes, reversible, confident → just ship.