The most expensive sentence in business
"We shipped it, and the numbers went up, so it worked." That sentence has justified more bad budget decisions than any spreadsheet error. It sounds like evidence. It is a story - one that quietly assumes the numbers would have stayed flat if you had done nothing. This session is about the missing half of that sentence: the world where you did not ship it. Get comfortable with that missing world - the counterfactual - and you can tell a genuine result from a lucky one, and know which decisions deserve the cost of a real test.
Causation is a comparison to a world you cannot see 8 min live
When Lumen's team says the new product page "caused" a lift, they are making a claim about two worlds: the one where these shoppers saw the new page, and the one where the very same shoppers, same day, saw the old one. Only one of those worlds ever happens. The gap between them is the true effect - and the entire craft of experimentation is a set of honest tricks for estimating a world you never got to observe.
LiveCorrelation, causation, and the confounder3 min▶
Ice-cream sales and drowning rise together. Neither causes the other - summer heat drives both. That hidden common cause is a confounder, and it is why "X and Y move together" almost never means "X causes Y".
- Correlation - two things move together. Cheap, everywhere, often coincidental.
- Causation - changing one changes the other. Expensive to prove, and the only thing worth betting budget on.
- Confounder - a third factor driving both, manufacturing a correlation with no causal link.
Lumen finds its email subscribers convert three times better than non-subscribers and nearly doubles the email budget. But subscribers chose to subscribe - they were already high-intent fans. The 3x is mostly who they are, not what email did. Engagement is the confounder. A holdout test later showed email's true incremental lift was a fraction of that. This is the single most common way marketing budgets get misallocated.
LiveWhy smart, honest people still get fooled4 min▶
It is not stupidity - it is that the wrong methods feel exactly as convincing as the right ones. Three traps catch everyone:
- The trend trap - attributing a seasonal or promo-driven rise to your change (the before/after picture above).
- The self-selection trap - comparing people who opted in to people who did not (the email story).
- The peeking trap - watching a live test daily and declaring victory the first day it looks good. The tool below shows how badly this one lies - and it feels like diligence, not error.
You do not need to run the statistics to defend against these. You need to recognize the shape of each and ask for the control group, the randomization, and the pre-committed stopping point.
Watching a test every day manufactures false wins try it
Here is a test where the two versions are genuinely identical - there is no real difference at all. Run it once checking only at the end, then run it checking every day and stopping the first time it looks significant. The "checking every day" version declares a false winner far more than the honest 5% of the time. Nobody lied; they just looked too often. This is why a good experimentation culture pre-commits to a stopping point.
When is a decision worth an experiment? 6 min live
Experiments cost traffic, time, and focus. Not every decision earns one. Your job as a leader is not to test everything - it is to spend your experimentation budget where the stakes and the uncertainty are both high, and to move fast on the rest. Here is the map.
Self-studyThe three questions to ask in any results review3 min read▶
You will spend more time reading experiments than commissioning them. These three questions, asked out loud, catch most bad claims - and you will have the vocabulary for all three by the end of the leader track.
- "What is the control group?" - if there is no comparable group that did not get the change, it is a before/after story, not a test. (Session 1, today.)
- "Was it powered, and did we pre-commit to the sample size?" - underpowered tests and peeked-at tests both produce fake wins. (Session 2 and 3.)
- "Did anything about the split look off - lopsided groups, a weird week?" - the trust checklist that separates a clean test from a broken one. (Session 3.)
This week ◐ 30 min total
- Find one "it worked" claim in a recent deck or Slack thread in your org. Ask yourself: was there a control group, or is it before/after? Write one sentence on what the missing counterfactual is.
- Play with the peeking tool above. Set the peeks high and low. Notice the false-win rate climb with more peeking even though nothing is truly different. Bring the intuition to Session 2.
- Map three upcoming decisions onto the stakes x uncertainty grid. Which genuinely deserve an experiment, and which are you overthinking?
- Optional: list your team's top guardrail metrics - the numbers that must NOT get worse even if the headline metric wins. We formalize these in Session 6.
Three questions before you go 🎯 ◐ 90 seconds
1 · A conversion chart rises the week after a launch. Why is that not proof the launch caused it?
Before/after has no control group, so seasonality, promos, and other shifts are baked into the "lift". The control group is your estimate of the counterfactual - what would have happened anyway.
2 · Email subscribers convert 3x better than everyone else. The safest reading is...
Self-selection makes the groups non-comparable. Only a holdout/randomized test isolates email's true incremental effect, which is usually far smaller than the raw gap.
3 · Which decision most clearly deserves a full experiment?
Experimentation budget belongs where stakes AND uncertainty are both high. Low-stakes, easy-to-reverse, confident changes should just ship - testing them costs more than being wrong.
What this session covers
This leader session distills the "why" that the technical courses assume you already believe. It covers ~80% of the conceptual foundation; the mechanics live in the builder track and the sources below.