Learn Experimentation with Phoebe

Did it actually work? Experimentation, A/B testing and causal inference, both halves.

Two tracks, one running brand. The 6-session leader track teaches executives, PMs, and marketers the thinking moves - read a results deck, spot an underpowered or peeked-at test, know when a geo test beats a button test, and build a culture that experiments instead of guessing. The 10-session builder track takes practitioners from potential outcomes and power analysis through A/B and multivariate testing, sequential testing and the peeking problem, CUPED variance reduction, geo experiments and synthetic control, difference-in-differences, and observational causal inference - designed and analyzed from scratch in Python. Built from the leading experimentation and causal-inference courses, and honest that the modern online-experiment stack (sequential, CUPED, synthetic control) is the gap no single course teaches.

16sessions, 2 tracks
9methods, from A/B to causal
2live in-browser simulators
1running brand: Lumen

The leader track ๐Ÿค ยท for execs, PMs & marketers ยท 6 x 45 min ยท no code

An experiment is only as trustworthy as the questions you ask of it. This track teaches you to read the number, spot the traps that fool smart people, and decide which decisions are even worth a test - without opening a notebook. A friendly on-ramp for LinkedIn readers, too.

The builder track ๐Ÿ› ๏ธ ยท for practitioners ยท 10 x 45 min ยท Python

For DA, DE, DS, and growth engineers. One running brand - Lumen Skincare - tested every way: from the potential-outcomes foundation and power analysis, through A/B and multivariate tests, the peeking problem and CUPED, to geo experiments, difference-in-differences, and honest observational causal inference - built from scratch on the same data.

๐Ÿ”ฎ Builder session 1 ยท start here

Potential outcomes and the estimand

The Rubin framework in code - two potential outcomes, one observed - and why one line (a coin flip) turns a biased guess into a causal answer.

โ–ถ Start herePython
๐Ÿ”‹ Builder session 2

Power and sample size

Type I and II error, power, and MDE - then compute the ~32,000 users per arm Lumen's test needs, with a live planner you rebuild in statsmodels.

Design the test
๐ŸŽฒ Builder session 3

Randomization and trust

Unit of randomization, SUTVA, stratification, the A/A test, and the sample-ratio-mismatch chi-square check that halts a broken experiment.

Protect validity
๐Ÿงช Builder session 4

Run and analyze an A/B test

The two-proportion z-test, confidence intervals, effect size, and the multiple-comparisons correction that makes a naive "winner" evaporate.

The payoff test
๐Ÿงฌ Builder session 5

Multivariate testing

Full vs fractional factorial, interaction effects, and two-way ANOVA - plus why aliasing hides the very interactions you hoped to find.

Many factors
๐Ÿ‘€ Builder session 6

Sequential testing and peeking

Why peeking inflates false positives to 30%+, and the fixes - always-valid inference, group-sequential boundaries, and bandits. With a live simulator.

Stop safely
๐Ÿ“‰ Builder session 7

CUPED and variance reduction

Use pre-experiment data to cut variance ~35% with zero bias - the Microsoft technique that shrinks Lumen's test from 32k to ~20k per arm.

Fewer users
๐Ÿ—บ๏ธ Builder session 8

Geo experiments and synthetic control

When you can only split by region - build a synthetic counterfactual from a weighted donor pool, validate the pre-period fit, and run placebo tests.

Region-level
๐Ÿ“Š Builder session 9

Difference-in-differences

Parallel trends, two-way fixed effects, and the event study - plus the staggered-adoption critique that broke the classic estimator.

Change-in-change
๐Ÿงฉ Builder session 10

Observational causal inference

Confounding, backdoor paths and DAGs, propensity scores and IPW, doubly-robust estimators - and the honest ceiling when you cannot randomize. Capstone.

Capstone
start here - foundations building up - the craft getting real - advanced methods hardest - causal & quasi-experimental

Choose your path ๐Ÿ—บ๏ธ

Leaders start at Leader session 1; practitioners start at Builder session 1. The fast-tracks fit a specific gap.

๐Ÿค Leader journey (no code) A1โ†’ A2โ†’ A3โ†’ A4โ†’ A5โ†’ A6โ†’ curious? Builder session 1
๐Ÿ› ๏ธ Full builder journey 1โ†’ 2โ†’ 3โ†’ 4โ†’ 5โ†’ 6โ†’ 7โ†’ 8โ†’ 9โ†’ 10
โšก Ship-a-clean-A/B fast-track 1โ†’ 2โ†’ 3โ†’ 4โ†’ 6
๐Ÿงฉ Causal-inference fast-track 1โ†’ 8โ†’ 9โ†’ 10
๐Ÿ‘” Leader crash course (60 min) A1โ†’ A3โ†’ A6
One brand, both tracks. Every session works on Lumen Skincare - a synthetic $18M direct-to-consumer skincare label with nine marketing channels and a $4M media budget. Its flagship test: does a new product-page layout lift checkout from 3.2% to 3.6%? Leaders learn to read and trust that result; builders design it (the ~32,000-user-per-arm calculation), run it, protect it from peeking, sharpen it with CUPED, and - when a clean A/B is impossible - reach for geo tests, difference-in-differences, and propensity scores on the very same data.
Honest about scope. Sessions distill the leading material - Udacity's A/B Testing (Google), 365 Data Science's A/B Testing in Python, Coursera's Crash Course in Causality (UPenn) and Experimentation for Improvement (McMaster), Brady Neal's causal inference course, and the key papers (Johari et al. on peeking, Deng-Xu-Kohavi on CUPED, Abadie on synthetic control, Card-Krueger and Callaway-Sant'Anna on DiD, Rosenbaum-Rubin on propensity). Certificates, videos, and graded assignments stay with their official sources. Content verified July 2026; re-check the fast-moving causal libraries (CausalImpact, DoWhy, EconML) before delivery.

The knowledge map ๐Ÿง 

The builder track at a glance - hover a session to spotlight its concepts, click any node to jump in.