learn-experimentation-with-phoebe / Leader session 2 of 6
Learn Experimentation with Phoebe · Leader session 2 of 6

Reading a test you didn't run

A results deck lands in your inbox. It has a green "significant" badge, a lift number, and a confident recommendation. Your job is not to redo the math - it is to know whether the math means what the slide says it means. This session teaches the four words that show up on every results slide, what "not statistically significant" does and does not tell you, and the two or three tells of a test that was too small or read too hopefully. No math, no code - just the reading comprehension that stops a fragile result from becoming a shipped decision.

🟢 Core Leaders · C-level / PM / marketing No code 45 min
0-3 · Setup 3-20 · The four words 20-34 · What p and CI claim 34-45 · Feel the tradeoff + wrap
Part 0

You are the last line of defence

By the time a test reaches your review, the data is collected and the analyst has done their honest best. The one thing still up for grabs is interpretation - and that is where good results turn into bad decisions. A slide that says "variant won, +12.5%, significant" can be true, or it can be a coin that happened to land heads. You cannot tell the difference by staring harder at the lift number. You tell the difference by reading four other numbers, and by knowing what silence (a "not significant" result) actually means. That is this session.

Live - covered in session Self-study - read after ★ The one big idea The case: Lumen Skincare
★ What you walk out with today A plain-language grip on the four words - significance, confidence interval, power, and MDE - so you can read any results deck without a statistician at your elbow. You will know why "not significant" is not the same as "no difference", why a result can be flagged "significant" and still probably be false, and how to spot the underpowered test that was doomed to be inconclusive before it ever ran. The confidence interval will become your favourite line on the slide.
Part 1 · vocabulary

The four words on every results slide 8 min live

Lumen ran its flagship test: the old product page converts at 3.2%, and the team hoped a "hero-first" layout would push it to 3.6%. That is the story on every slide in this session. Four words decide whether that 0.4 point jump is real. Here is each one in a single exec sentence, grounded in that exact test.

The Lumen test in four words: 3.2% control vs 3.6% variant, ~32,000 users per arm MDE Smallest lift you designed to catch. Here: +0.4pt (a +12.5% relative). Power Odds of catching a real lift that size. Here: 80% - so a 1-in-5 chance you miss. Significance How surprised you'd be if nothing were going on. p<0.05 = "quite surprised". Confidence interval The honest range the true lift likely sits in. Narrow = precise. Read this one. MDE and power are decided BEFORE the test (they size it). Significance and the confidence interval are read AFTER (they judge it). A deck that shows a lift but hides the confidence interval is telling you the headline and hiding the uncertainty. If you remember one thing: ask for the confidence interval. It contains the lift, the precision, and the honesty in one line.
🔍 Click to zoom - the four words, sized before the test and read after it
LiveMDE and power - the two you set before the test4 min

These two words are design choices, not findings. They are agreed before a single visitor is bucketed, and they quietly decide whether the test can succeed at all.

  • MDE (minimum detectable effect) - the smallest lift you promise to be able to see. For Lumen it is +0.4 points (3.2% to 3.6%, a +12.5% relative lift). Set the MDE smaller and you need far more traffic; set it larger and the test is cheap but blind to modest wins.
  • Power - if a real lift of that size exists, the chance the test actually flags it. The industry default is 80%, which means a fully correct test still misses a true winner one time in five. Power below 80% is a test that is likely to come back "nothing here" even when something is there.
Real world

Lumen's analyst sized the test for a +0.4 point MDE at 80% power and 95% significance. That is why the plan called for roughly 32,000 users per arm, about 64,000 in total, running around 16 to 17 days at 4,000 checkout sessions a day. Those are not arbitrary - they are the price of being able to detect a lift that small with confidence. If a rushed stakeholder had said "let's just run it for three days", the test would have been underpowered and almost guaranteed to be inconclusive, wasting the traffic entirely.

LiveSignificance and the confidence interval - the two you read after4 min

These two are the results. They arrive when the test is done, and they are what you actually interrogate in the review.

  • Statistical significance (the p-value) - a measure of surprise. It answers: if the two pages were truly identical, how often would pure luck hand us a gap this big or bigger? A p-value below 0.05 means "less than 1 time in 20" - surprising enough that we stop calling it luck. It is a threshold, not a truth.
  • Confidence interval - the range of lifts consistent with the data, usually written as "95% CI". Instead of a single number it gives you a band, for example "+0.4 points, 95% CI +0.1 to +0.7". The width tells you the precision; whether it crosses zero tells you if the result is conclusive.
Question to ask in a review "What is the confidence interval, not just the point estimate?" A single lift number always looks precise. The interval reveals whether the true effect could be a strong win, a wash, or even a small loss - which is the difference between a decision and a guess.
Part 2 · the traps in the words

What "significant" claims, and what it quietly does not 6 min live

The single most misread phrase in any deck is "not statistically significant". People hear it as "we proved there is no difference". It means almost the opposite: we did not gather enough evidence to be sure there is one. Absence of evidence is not evidence of absence. And even a "significant" result carries a trap of its own when a team tests many things at once.

Live"Not significant" means "not proven", not "no effect"3 min

A "not significant" result has two very different possible causes, and the slide usually will not tell you which:

  • There genuinely is no meaningful effect - the change did nothing.
  • There is an effect, but the test was too small to see it - underpowered, so the noise swamped the signal.

These look identical on the slide and mean opposite things for your decision. The confidence interval is what tells them apart: a tight interval hugging zero really is evidence of "no meaningful effect", while a wide interval that happens to include zero is just an inconclusive test that could not decide.

Real world

Lumen tested a free-shipping banner and the deck read "no significant lift, do not ship". But the confidence interval ran from -0.5 to +1.3 points - enormously wide, straddling zero. That was not proof the banner did nothing; it was proof the test was too small to tell. Reading "not significant" as "proven useless", the team almost killed an idea that a properly powered rerun later showed was a genuine win. "Not significant" was the analyst being honest about ignorance, not a verdict.

LiveThe base-rate trap - "significant" can still be mostly wrong3 min

Here is the counter-intuitive one. If a team runs many tests, most of which are long shots, then a good share of the "significant" wins they celebrate will be flukes - even with the 5% threshold working exactly as designed.

  • Say a team tries 20 ideas and, realistically, only 2 of them truly work.
  • Of the 18 duds, roughly 1 will pass the 5% significance bar by pure luck (that is what 5% means).
  • So the winners board reads "3 significant results" - but 1 of the 3 is a mirage. A third of your "wins" are noise, despite doing the statistics correctly.

The fix is not to distrust statistics - it is to raise the prior. Test ideas that have a real reason to work, replicate surprising wins before betting big on them, and be extra sceptical of a "significant" result that nobody predicted.

Question to ask in a review "How many things did we test to find this winner?" A single win pulled from 30 quiet experiments deserves a confirmation run before it reshapes the roadmap.
Part 3 · your favourite line on the slide

The confidence interval is the honest summary 5 min live

If you only trained yourself to read one thing, make it the confidence interval. It folds three separate judgements into a single line: how big the effect is, how precisely you measured it, and whether the result is even conclusive. The two questions to ask of any interval are "how wide is it?" (precision) and "does it cross zero?" (conclusiveness).

no effect (zero) ← worse better → Clear win +0.1 to +0.7pt - entire range is positive. Ship. Inconclusive -0.5 to +1.3pt - straddles zero. Not proven either way. Clear loss -0.7 to -0.1pt - entire range is negative. Do not ship. The rule of thumb: if the whole interval is on one side of zero, you have a direction. If it crosses zero, the test could not decide - do not read it as "no effect".
🔍 Click to zoom - does the interval cross zero? that one question decides conclusiveness
zero true lift +0.4pt Small sample ~3,000 / arm Interval so wide it swallows zero AND the true lift. Comes back "not significant" - but proves nothing. Right sample ~32,000 / arm Narrow interval sits clear of zero. Same true lift, now visible - a real win. Underpowered is not "unlucky" - it is a test that was mathematically unable to give a clear answer before it started.
🔍 Click to zoom - too small a sample hides a real effect inside a wide interval
Self-studyThe tells of a test that cannot be trusted at face value3 min read

You do not need to recompute anything to smell an over-interpreted test. Watch for these on the slide:

  • A very wide confidence interval - the sample was too small; the headline lift is a guess dressed as a finding.
  • No confidence interval at all - only a point estimate and a green badge. The uncertainty was left off the slide.
  • A tiny sample with a huge claimed lift - big effects from small tests usually shrink or vanish on a rerun (the "winner's curse").
  • "Not significant" read as "no effect" - ask whether the test was powered to detect the effect size that would matter to the business.
  • A short run that stopped early on a good day - which is the peeking trap you met in Session 1, and the subject of the checklist in Session 3.
Question to ask in a review "Was this test powered to detect the smallest lift we'd actually act on - and if it came back flat, is the interval tight enough to call it a real 'no'?" That single question separates a confident decision from wishful reading.
Feel the sample-size tradeoff

Why detecting a small lift costs so much traffic try it

This is the exact Lumen test. The baseline is 3.2%, and you are trying to detect a +12.5% relative lift (the jump to 3.6%). Drag the dials and watch the sample size move. Two lessons live in this tool: chasing a smaller lift explodes the traffic you need, and demanding more power (fewer missed winners) costs traffic too. This is why analysts pin down the MDE and power before a test - they are quietly setting its price.

Before Session 3

This week ◐ 30 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · A deck says "the new page showed no statistically significant lift." What is the safest reading?

Absence of evidence is not evidence of absence. Check the confidence interval: a tight band near zero is a real "no"; a wide band that crosses zero is an inconclusive, underpowered test.

2 · A results slide shows "+0.6pt lift, 95% CI -0.2 to +1.4." How should you act?

When the confidence interval straddles zero, the test could not decide direction. The point estimate looks like a win, but the honest range includes "no effect" and even a small loss.

3 · Lumen's flagship test needs ~32,000 users per arm. Why can't the team just run it for three days?

The +0.4pt MDE at 80% power requires roughly 64,000 users total, about 16-17 days at 4,000/day. A three-day run has too little traffic to detect a lift that small, so it is doomed to be inconclusive before it starts.

Source material

What this session covers

This leader session teaches the reading comprehension for a results deck - the conceptual ~80% of the statistics without the computation. The mechanics (how the numbers are actually calculated) live in Builder Session 2 and the sources below.

MDE, sample size, and power in plain languageexec framing - Part 1
p-value / significance and the base-rate trapwhat "significant" claims - Part 2
Confidence intervals, "not significant" ≠ "no effect", underpowered tellsPart 2 and Part 3
365 Data Science - A/B Testing in Python (Kuznetsova)the on-point mechanics - Builder Session 2
Google - The Power of Statistics (GADA C3)hypothesis testing and CIs in depth, no math needed here
The rule-of-16 and two-proportion sample-size mathBuilder Session 2 - how ~32,000/arm is derived

Leader Session 2 cheat sheet · pin this

MDESmallest lift you designed to catch. Lumen: +0.4pt (+12.5% relative). A design input, set before the test.
PowerOdds of catching a real lift that size. 80% is standard - a correct test still misses a true winner 1 time in 5.
Significance (p<0.05)A measure of surprise: less than 1-in-20 that pure luck produced a gap this big. A threshold, not a truth.
Confidence intervalThe honest range for the true effect. Width = precision; crossing zero = inconclusive. Always ask for it.
"Not significant" ≠ "no effect"Absence of evidence is not evidence of absence. A wide interval means the test was too small to decide.
Base-rate trapTest many long shots and some "significant" wins are flukes. Ask how many things were tested; replicate surprises.