You are the last line of defence
By the time a test reaches your review, the data is collected and the analyst has done their honest best. The one thing still up for grabs is interpretation - and that is where good results turn into bad decisions. A slide that says "variant won, +12.5%, significant" can be true, or it can be a coin that happened to land heads. You cannot tell the difference by staring harder at the lift number. You tell the difference by reading four other numbers, and by knowing what silence (a "not significant" result) actually means. That is this session.
The four words on every results slide 8 min live
Lumen ran its flagship test: the old product page converts at 3.2%, and the team hoped a "hero-first" layout would push it to 3.6%. That is the story on every slide in this session. Four words decide whether that 0.4 point jump is real. Here is each one in a single exec sentence, grounded in that exact test.
LiveMDE and power - the two you set before the test4 min▶
These two words are design choices, not findings. They are agreed before a single visitor is bucketed, and they quietly decide whether the test can succeed at all.
- MDE (minimum detectable effect) - the smallest lift you promise to be able to see. For Lumen it is +0.4 points (3.2% to 3.6%, a +12.5% relative lift). Set the MDE smaller and you need far more traffic; set it larger and the test is cheap but blind to modest wins.
- Power - if a real lift of that size exists, the chance the test actually flags it. The industry default is 80%, which means a fully correct test still misses a true winner one time in five. Power below 80% is a test that is likely to come back "nothing here" even when something is there.
Lumen's analyst sized the test for a +0.4 point MDE at 80% power and 95% significance. That is why the plan called for roughly 32,000 users per arm, about 64,000 in total, running around 16 to 17 days at 4,000 checkout sessions a day. Those are not arbitrary - they are the price of being able to detect a lift that small with confidence. If a rushed stakeholder had said "let's just run it for three days", the test would have been underpowered and almost guaranteed to be inconclusive, wasting the traffic entirely.
LiveSignificance and the confidence interval - the two you read after4 min▶
These two are the results. They arrive when the test is done, and they are what you actually interrogate in the review.
- Statistical significance (the p-value) - a measure of surprise. It answers: if the two pages were truly identical, how often would pure luck hand us a gap this big or bigger? A p-value below 0.05 means "less than 1 time in 20" - surprising enough that we stop calling it luck. It is a threshold, not a truth.
- Confidence interval - the range of lifts consistent with the data, usually written as "95% CI". Instead of a single number it gives you a band, for example "+0.4 points, 95% CI +0.1 to +0.7". The width tells you the precision; whether it crosses zero tells you if the result is conclusive.
What "significant" claims, and what it quietly does not 6 min live
The single most misread phrase in any deck is "not statistically significant". People hear it as "we proved there is no difference". It means almost the opposite: we did not gather enough evidence to be sure there is one. Absence of evidence is not evidence of absence. And even a "significant" result carries a trap of its own when a team tests many things at once.
Live"Not significant" means "not proven", not "no effect"3 min▶
A "not significant" result has two very different possible causes, and the slide usually will not tell you which:
- There genuinely is no meaningful effect - the change did nothing.
- There is an effect, but the test was too small to see it - underpowered, so the noise swamped the signal.
These look identical on the slide and mean opposite things for your decision. The confidence interval is what tells them apart: a tight interval hugging zero really is evidence of "no meaningful effect", while a wide interval that happens to include zero is just an inconclusive test that could not decide.
Lumen tested a free-shipping banner and the deck read "no significant lift, do not ship". But the confidence interval ran from -0.5 to +1.3 points - enormously wide, straddling zero. That was not proof the banner did nothing; it was proof the test was too small to tell. Reading "not significant" as "proven useless", the team almost killed an idea that a properly powered rerun later showed was a genuine win. "Not significant" was the analyst being honest about ignorance, not a verdict.
LiveThe base-rate trap - "significant" can still be mostly wrong3 min▶
Here is the counter-intuitive one. If a team runs many tests, most of which are long shots, then a good share of the "significant" wins they celebrate will be flukes - even with the 5% threshold working exactly as designed.
- Say a team tries 20 ideas and, realistically, only 2 of them truly work.
- Of the 18 duds, roughly 1 will pass the 5% significance bar by pure luck (that is what 5% means).
- So the winners board reads "3 significant results" - but 1 of the 3 is a mirage. A third of your "wins" are noise, despite doing the statistics correctly.
The fix is not to distrust statistics - it is to raise the prior. Test ideas that have a real reason to work, replicate surprising wins before betting big on them, and be extra sceptical of a "significant" result that nobody predicted.
The confidence interval is the honest summary 5 min live
If you only trained yourself to read one thing, make it the confidence interval. It folds three separate judgements into a single line: how big the effect is, how precisely you measured it, and whether the result is even conclusive. The two questions to ask of any interval are "how wide is it?" (precision) and "does it cross zero?" (conclusiveness).
Self-studyThe tells of a test that cannot be trusted at face value3 min read▶
You do not need to recompute anything to smell an over-interpreted test. Watch for these on the slide:
- A very wide confidence interval - the sample was too small; the headline lift is a guess dressed as a finding.
- No confidence interval at all - only a point estimate and a green badge. The uncertainty was left off the slide.
- A tiny sample with a huge claimed lift - big effects from small tests usually shrink or vanish on a rerun (the "winner's curse").
- "Not significant" read as "no effect" - ask whether the test was powered to detect the effect size that would matter to the business.
- A short run that stopped early on a good day - which is the peeking trap you met in Session 1, and the subject of the checklist in Session 3.
Why detecting a small lift costs so much traffic try it
This is the exact Lumen test. The baseline is 3.2%, and you are trying to detect a +12.5% relative lift (the jump to 3.6%). Drag the dials and watch the sample size move. Two lessons live in this tool: chasing a smaller lift explodes the traffic you need, and demanding more power (fewer missed winners) costs traffic too. This is why analysts pin down the MDE and power before a test - they are quietly setting its price.
This week ◐ 30 min total
- Find a real results deck in your org - any A/B test, experiment, or "we measured the impact" slide. Restate its headline claim as a confidence interval in one sentence: "the true effect is somewhere between X and Y." If the deck does not give you enough to do that, that itself is the finding.
- Spot whether it was powered. Was an MDE and power agreed before the test, or did the team run it "until it looked done"? Write one line on which.
- Play with the planner above. Halve the MDE and watch the sample size climb. Push power from 80% to 95% and see the cost. Bring the intuition that "detecting small lifts is expensive" to Session 3.
- Optional: find one deck that says "no significant difference" and ask whether the interval was tight (a real "no") or wide (an inconclusive test wearing a verdict).
Three questions before you go 🎯 ◐ 90 seconds
1 · A deck says "the new page showed no statistically significant lift." What is the safest reading?
Absence of evidence is not evidence of absence. Check the confidence interval: a tight band near zero is a real "no"; a wide band that crosses zero is an inconclusive, underpowered test.
2 · A results slide shows "+0.6pt lift, 95% CI -0.2 to +1.4." How should you act?
When the confidence interval straddles zero, the test could not decide direction. The point estimate looks like a win, but the honest range includes "no effect" and even a small loss.
3 · Lumen's flagship test needs ~32,000 users per arm. Why can't the team just run it for three days?
The +0.4pt MDE at 80% power requires roughly 64,000 users total, about 16-17 days at 4,000/day. A three-day run has too little traffic to detect a lift that small, so it is doomed to be inconclusive before it starts.
What this session covers
This leader session teaches the reading comprehension for a results deck - the conceptual ~80% of the statistics without the computation. The mechanics (how the numbers are actually calculated) live in Builder Session 2 and the sources below.