learn-experimentation-with-phoebe / Leader session 6 of 6
Learn Experimentation with Phoebe · Leader session 6 of 6 · capstone

Building an experimentation culture, the leader's job

You have learned to read a test, trust a test, and know which test fits which question. The last thing a leader owns is the system around the tests - the culture that decides what "winning" means before a test runs, protects the things that must not break, pre-commits to a decision, and celebrates good tests rather than only good results. This capstone hands you the four levers and the anti-patterns to watch for, then puts you in the chair: commission Lumen's experimentation roadmap. No math, no code.

🟡 Core Leaders · C-level / PM / marketing No code 45 min
0-3 · Setup 3-24 · OEC, guardrails, decisions 24-40 · Velocity & anti-patterns 40-45 · Capstone
Part 0

Good tests are a habit, not an event

A single well-run experiment is worth something. An organization that runs a hundred a year, each one honest, each one feeding the next decision, is worth a great deal more - and that is a leadership artifact, not a statistical one. The statistics are settled by Sessions 1 to 5. What is left is the part only you can build: agreeing on the success metric up front, protecting the numbers that must not break, deciding in advance what a result will mean, and building a team that treats a failed test as cheap learning rather than a black mark. This capstone is about that system.

Live - covered in session Self-study - read after ★ The one big idea The case: Lumen Skincare
★ What you walk out with today Four levers you can pull on Monday: the OEC (one agreed success metric per test, chosen before it runs), guardrail metrics (the things that must not get worse), a pre-committed decision framework (ship / iterate / kill), and a view of experiment velocity and ROI that rewards good tests over lucky wins. Plus the org anti-patterns that quietly rot all four. You finish by commissioning Lumen's roadmap, closing the loop back to Session 1's stakes-and-uncertainty grid.
Part 1 · the one big idea

Decide what winning means before the test runs 9 min live

The most common way a test goes wrong has nothing to do with statistics. It is that nobody agreed, before the test, what result would count as a win - so afterwards everyone finds the metric that flatters their side. The fix is to name one Overall Evaluation Criterion (OEC) up front: the single number this test is judged on. Then name the guardrails - the metrics that must not get worse even if the OEC wins. And then pre-commit to what each outcome triggers: ship, iterate, or kill.

The OEC: one headline metric this test is judged on Checkout conversion rate (OEC) 3.2% → 3.6% ✓ OEC wins Guardrails: must not get worse Page-load latency ✗ FAIL +180ms slower Refund rate ✓ PASS flat Add-to-cart rate ✓ PASS +0.2pt Unsubscribe rate ✓ PASS flat Verdict: OEC won but a guardrail failed → do NOT ship as-is. Fix the latency regression, then re-test. A win that breaks a guardrail is not a win.
🔍 Click to zoom - one headline OEC over a row of guardrails; a broken guardrail vetoes the win
LiveThe OEC and why it must be singular3 min

The Overall Evaluation Criterion is the one metric a test is graded on, agreed before launch. Its whole power is that it is singular and pre-committed - it takes away the after-the-fact temptation to cherry-pick whichever number moved.

  • One metric, chosen up front - for Lumen's page test, checkout conversion. Not "conversion, or revenue, or time-on-page, whichever looks good later".
  • It should capture long-term value, not a vanity spike - clicks can rise while revenue falls. A good OEC is the number you would actually stake the business on.
  • Agreed by the people who will act on it - so the result is binding, not debatable.
Real world

Lumen's hero-first page test names checkout conversion as its OEC before launch - target lift from 3.2% to 3.6%. When the test comes in at exactly that, there is no argument about whether it "worked", because the bar was set in advance. That is the OEC doing its job.

LiveGuardrails and the pre-committed decision4 min

Guardrail metrics are the things that must not get worse even when the OEC wins. They stop a team from buying a headline number at the cost of something quietly expensive.

  • Lumen's guardrails - page-load latency, refund rate, add-to-cart rate, unsubscribe rate. A conversion win that slows the page by 180ms is not shipped until the latency is fixed.
  • The decision, pre-committed - before the test, agree what each outcome triggers: ship (OEC wins, guardrails hold), iterate (promising but not conclusive, or a guardrail wobbled), kill (flat or negative). Deciding in advance removes the emotion afterwards.
  • Absence of evidence is not evidence of absence - a "not significant" result from Session 2 usually means iterate or re-power, not "the idea is dead".
The one sentence to institutionalize "What is the OEC, what are the guardrails, and what will each result make us do?" Ask it before every test and you have installed most of an experimentation culture with a single question.
Part 2 · the system

Velocity, the ROI of failure, and what rots a culture 7 min live

Most experiments fail - the change does not move the OEC, or it moves it the wrong way. In a healthy culture that is the point, not a problem. Each cheap failed test buys a piece of knowledge and steers the next bet. The return on experimentation comes from volume plus honesty: run many tests, learn from all of them, ship the few real winners. The leader's job is to make failing safe, keep the loop turning, and stamp out the behaviours that quietly turn the whole system into theatre.

Hypothesis Design Run Decide Learn Repeat Most loops end in "kill" - and that is cheap learning that steers the next bet. Velocity is the flywheel: the faster the loop turns and the more honestly it reports, the more the whole org learns. Reward good tests, not just wins.
🔍 Click to zoom - hypothesis to design to run to decide to learn, and around again
LiveThe ROI of experimentation is mostly learning3 min

If you only counted the winning tests, experimentation would look inefficient - most tests do not win. That framing misses the point. The value is in the portfolio: cheap failures that stop you shipping expensive mistakes, plus the occasional real winner that pays for all of them.

  • Most tests fail, by design - if nearly all your tests win, you are testing only safe bets and learning little. A healthy program has plenty of flat and negative results.
  • Failures are cheap insurance - a $32,000-per-arm test that kills a bad idea saved you from rolling that idea out to everyone.
  • Reward the test, not the outcome - a well-designed test that returned a clean "no" is a success. Celebrating only positive results teaches teams to game the metric.
Real world

Lumen's growth team runs about 40 tests a year. Roughly 8 ship, a handful iterate, the rest are killed. The CMO reports the program by learnings banked and mistakes avoided, not by win rate - which is exactly why the team keeps proposing bold, risky tests instead of safe ones.

Self-studyThe org anti-patterns that rot experimentation3 min read

Every failure mode here is cultural, not statistical - which means only a leader can fix it.

  • The HiPPO override - the Highest-Paid Person's Opinion overturns a clean result because they do not like it. Kills the incentive to test honestly. If you commission an experiment, you agree to be bound by it.
  • P-hacking to a win - slicing the data until some segment looks significant, then reporting that. The peeking trap from Session 1, dressed up. Pre-commit the OEC and the analysis.
  • Celebrating only positive results - if wins get applause and honest "no" results get silence, teams learn to manufacture wins. Reward the quality of the test.
  • No pre-committed decision - running a test with no agreement on what the outcome triggers, so the result gets argued away.
The leader's standing rule The person who commissions the test does not get to veto it afterwards. That single commitment defends against most of these at once.
Capstone · put it together

Commission Lumen's experimentation roadmap ◐ 30 min

This is the whole leader track in one exercise. Lumen's CMO must reallocate a $4M media budget defensibly across 9 channels. Play the leader who commissions the evidence.

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · Why name the OEC before the test runs rather than after?

The OEC's power is that it is singular and agreed up front. Choosing the metric after seeing the data is how p-hacking and motivated reasoning creep in.

2 · Lumen's page test lifts conversion to target, but page-load latency gets 180ms worse. What should happen?

Guardrails are the metrics that must not get worse even when the OEC wins. A conversion win bought with a latency regression is not a win until the guardrail is restored.

3 · Most of Lumen's tests come back flat or negative. What does that tell a good leader?

Most experiments fail by design, and that is the point. The ROI comes from cheap failures that prevent expensive mistakes plus the occasional real winner. Reward good tests, not just wins.

Source material

What this session covers

This capstone distills the culture chapters of the experimentation literature - the part the technical courses assume a leader already owns. It covers ~80% of the conceptual content; the rest lives in the sources below.

The OEC - one agreed success metric per test, chosen up frontexec framing - Part 1
Guardrail metrics and the pre-committed ship / iterate / kill decisionthe dashboard concept - Part 1
Experiment velocity, the ROI of failure, org anti-patternsthe lifecycle loop - Part 2
Capstone - commission Lumen's roadmap from the $4M budgetcloses the loop to Session 1
Kohavi, Tang & Xu - Trustworthy Online Controlled Experimentsthe culture, OEC, and guardrail chapters - the industry bible
The full builder track (b1-b10)the statistics behind every tool you now know how to commission

Leader Session 6 cheat sheet · pin this

OECOverall Evaluation Criterion: the one metric a test is judged on, agreed before it runs. Singular and pre-committed on purpose.
Guardrail metricsThe numbers that must not get worse even if the OEC wins - latency, refunds, unsubscribes. A win that breaks one is not shipped.
Decision frameworkShip / iterate / kill, pre-committed before launch so the result is binding, not debatable.
Velocity & ROIMost tests fail, and that is the point. Value = cheap failures avoided + rare real winners. Reward good tests, not just wins.
Anti-patternsHiPPO overrides, p-hacking to a win, celebrating only positive results, no pre-committed decision. All cultural, all yours to fix.
The standing ruleWhoever commissions the test does not get to veto it afterwards. One commitment that defends against most anti-patterns at once.