learn-ai-pm-with-phoebe / PM session 8 of 10
Learn AI Product Management with Phoebe · PM track · Session 8 of 10

Experiments, and the readout that says no

This session is not about statistics. It is about the two documents around the test: the design you write before any data exists, and the readout you write after. Both are decision documents, and both are where teams quietly cheat - the design by writing a hypothesis that cannot lose, the readout by explaining the result instead of reporting it. The worked example on this page is Cadence's summaries rollout, and it comes out mixed, because a course that only shows wins teaches nothing about readouts.

🔴 PM track PMs · founders · product leads Worked artifact · a real readout 45 min
0-3 · Welcome 3-16 · What the test is for 16-40 · Rules before data, then the readout 40-45 · Q&A
Part 0

Why the readout is the real skill

b7 set the numbers: week-1 transcript revisit rate as the primary metric, 11% today against a 30% target, self-reported in-meeting note-taking as the secondary at 68%, and accuracy complaints as the guardrail at or below 1.4% of meetings. The events are specified. Now something ships, and the organisation has to find out what happened without flattering itself.

Almost every team can design a test. Very few write a readout that is allowed to end in "we are stopping". That document is the one nobody is rewarded for, it is the one that saves the next quarter, and it is most of this session. The statistics side of experiments - power, variance reduction, sequential testing - is learn-experimentation, deliberately not re-taught here.

Live - presented in session Self-study - read after class ▶ Worked artifact - the readout, in full Sources covered
★ What you walk out with today A decision rule you wrote before the data existed, the three test shapes that beat an A/B test at your size, and a readout template that is structurally capable of saying no.
Part 1 · covers experiment practice: hypotheses, decision rules

An experiment is for deciding, not for proving 7 min live

Start with the cheapest test in this whole course, and it is not a test at all. Write down what you would do if the result came back positive, and what you would do if it came back negative. If the two answers are the same, do not run the experiment. You already decided, and what you are about to buy is not information, it is cover.

Proving · what most tests are secretly for the hypothesis is written so it cannot lose success is whatever the data supports you stop looking whenever it looks good the readout argues for what already shipped a null result becomes a learning story Tell: what result would have stopped us? Deciding · what a test is actually for the hypothesis names a number that can miss the decision rule is written before the data the stop is a date and a threshold, agreed the readout can end in stop, and sometimes does a null result is a result, and it is cheap Tell: which outcome changes what we do? Cadence's decision rule, written two weeks before launch and not edited since: "If week-1 transcript revisit rate has not reached 20% by the end of week 3, we stop the phase-2 digest build and spend two weeks on the six highest-revisit accounts instead. Summaries stay on either way."
🔍 Click to zoom - the same test, two postures, two completely different readouts
LiveWriting a hypothesis that is allowed to lose4 min

A hypothesis is falsifiable or it is a slogan. Compare the two versions of the same belief about Cadence:

  • Cannot lose: "Post-meeting summaries will improve the experience for team admins and drive engagement." Every observed outcome is consistent with this sentence. It cannot be wrong, so it cannot be informative, and the readout that follows it will be a paragraph of adjectives.
  • Can lose: "If team admins on paid workspaces of 5 to 50 seats get a summary within 90 seconds of a call ending, week-1 transcript revisit rate rises from 11% to 30% within four weeks, while accuracy complaints stay at or below 1.4% of meetings." Named segment, named number, named baseline, named horizon, named guardrail. It can miss, and this one did.

Then attach the decision rule in the same breath, in the same document, with a date on it. Not "we will review the results" - the actual if-then, including the branch where you stop. If you cannot write the stopping branch, you are not designing a test; you are scheduling a celebration.

The two-answer test, before anything else Write what you do if it wins and what you do if it loses. Same answer? Do not run it. Different answers? The difference between them is exactly the value of the experiment, and now you know what you are buying.
LiveThe decision rule, and where teams put it3 min

A decision rule has four parts, and a rule missing any one of them is a wish: the metric, the threshold, the date, and the action on each side. Cadence's rule in the figure above has all four, which is why it was able to fire without a meeting.

Put it somewhere a stakeholder will see it before the result exists. In the spec, in the launch note, in the channel where the launch is announced. A rule that lives only in your head is not a rule; it is a memory that will be revised by whatever the number turns out to be. The only real protection is that other people read the threshold while it was still uncomfortable.

And write the action on the losing side concretely enough that it can be executed by someone who is disappointed. "We stop the phase-2 digest build and spend two weeks on the six highest-revisit accounts" survives a bad Monday. "We will reconsider our approach" does not survive lunch.

Self-studyThree tests that should never have been run3 min read
  • The already-decided test. The build is committed, the launch date is public, and the test is running to produce a chart for the launch email. Cost: three weeks of instrumentation and the credibility of every future test.
  • The unpowered test on a small segment. A change tested on 40 accounts where only a very large effect would be visible, then reported as "no significant difference" and used to kill the idea. You did not learn it does not work. You learned nothing, expensively, and you learned it in a way that sounds like knowledge.
  • The test nobody can act on. The result arrives four weeks after the planning cycle closed. Whatever it says, the roadmap is set. Run the test that lands before the decision, or move the decision, or do not run it.

All three failures happen before any data exists, which is the pattern of this session: the expensive mistakes in experimentation are design mistakes and readout mistakes, not analysis mistakes.

Part 2 · covers stopping rules, minimum effect, test shapes

Three things you decide before a single event fires 6 min live

The order is the whole lesson. Set the stopping rule, the minimum effect worth shipping and the decision rule while the data does not exist yet, because once it exists you are no longer designing a test. You are negotiating with yourself, and you always win that negotiation.

Written down before launch, read by someone other than you 1 · The stopping rule When you stop looking: a date, a sample, or both. Decided now. Because after the data lands you will find a reason to wait. 2 · Minimum effect worth it The smallest move that would make you ship anyway. Here it is 11% to 20%, not 11% to 12%. Below it, shipping costs more. 3 · The decision rule Metric, threshold, date, and the action on each side. One sentence. Including the side where you stop. Any threshold set after the data exists gets set exactly where the data already is. That is not analysis, it is narration. The dates and numbers above are worth more than the analysis that follows them.
🔍 Click to zoom - three rules, all of them cheap, all of them worthless if written late
LiveFour shapes, and three of them are not an A/B test4 min

A team of Cadence's size, with a segment of paid workspaces of 5 to 50 seats, often cannot run a well-powered A/B test on a metric like week-1 revisit rate inside a useful window. That is not a reason to skip the discipline. It is a reason to pick a shape that fits, and to be honest in the readout about what the shape can and cannot tell you.

ShapeWhat it isUse it whenWhat it cannot tell you
A/B testrandomised split, one change, a measured differenceyou have the traffic, the change is isolated, and the decision can wait for the windowanything about accounts too rare to land in both arms
Staged rollout with a guardrail20% then 60% then all, with a pre-agreed number that pauses ityou are fairly confident, the risk is a regression rather than a dud, and speed mattersa clean causal estimate: the world moved while you rolled out
Painted doorthe entry point exists, the feature does not; you measure intent to enterthe build is expensive and you are unsure anyone wants the door at allwhether they would keep using it once the room behind the door is real
Concierge testten accounts, the outcome delivered by hand, no product builtyou need to know whether the outcome is valued before you automate itanything about scale, cost, or self-serve discoverability

Cadence shipped summaries as a staged rollout with a guardrail, and the readout below says so in its first paragraph. That single admission is what keeps the rest of the document honest, because it tells the reader which claims are causal and which are merely observed.

Self-studyThe concierge test nobody runs, and should3 min read

The CEO's agent ask has no evidence behind it. The instinct is to argue about it or to build a prototype, and both are expensive. The concierge shape is neither: pick ten workspaces, and for two weeks have a human do by hand whatever the agent would have done - chase the open action items, send the follow-up, update the tracker. No product. A shared doc and somebody's afternoon.

Two weeks later you know something nobody in the argument knew: whether admins let it happen without watching, what they corrected, and what they asked for that was not in the pitch. That is evidence, it cost no engineering, and it is exactly the discovery task b9 converts the CEO's ask into. Ten accounts and a human beats a quarter of building and a debate.

Part 3 · covers honest readouts

The five questions a readout answers 5 min live

A readout is not a summary of what happened. It is the document that converts a result into a decision, in front of witnesses, and its structure is the thing that makes honesty the path of least resistance. Answer these five in this order and the awkward parts have nowhere to hide.

Five sections, in this order, whatever the result turns out to be 1 · What we shipped The change, the population it reached, the shape of the test, the dates. 2 · What we predicted The numbers written before the data, unedited, including the wrong ones. 3 · What happened Primary, secondary, guardrail, cost. Each against its prediction. 4 · What we now believe The interpretation, plus the confounds you can name and cannot rule out. 5 · What we are doing The decision rule's consequence, executed. Not a plan invented today. A readout that cannot end in "we are stopping" is not a readout. It is a launch announcement with a chart in it, and everyone in the room can tell, including the people nodding.
🔍 Click to zoom - the order matters: prediction before result, result before interpretation
LiveWhy prediction comes before result3 min

Section 2 exists to make section 4 hard to fake. If the reader sees the prediction first, "8 points of lift in four weeks" cannot be presented as a triumph when the number written down was 19 points. Put the result first and the prediction becomes optional context, then it becomes a footnote, then it disappears, and the readout becomes a highlight reel with a methodology section.

The same logic makes section 5 a consequence rather than a proposal. If the decision rule is already in section 2, then the action in section 5 is arithmetic that anyone in the room can check. If it is not, section 5 is a new argument you are making while everybody is emotional about the number, which is the worst possible time to make one.

Say the mixed result in the first two lines Not at the end, not in an appendix. "Revisit rate improved by 8 points against a target of 19, note-taking did not move, the guardrail held, and the decision rule fired" is the honest opening sentence. Everything after it is detail, and nobody has to hunt.
Self-studyThe organisational problem behind bad readouts3 min read

People do not write dishonest readouts because they are dishonest. They write them because the last five honest ones were met with silence and the last five launch announcements were met with congratulations. The incentive is doing exactly what incentives do.

You cannot fix that alone, but you can do three things that measurably help. Write the decision rule publicly, so a mixed result reads as the system working rather than as your failure. Publish a readout for something that worked and include what you got wrong in it, so the format is not associated only with bad news. And when somebody else publishes a readout that ends in stop, say so out loud in the channel where the launch would have been celebrated. Three quarters of that and the format stops being career risk.

Worked artifact · 12 min · everyone

The actual readout for Cadence's summaries launch ★ mixed result, on purpose

This is the whole document, not an excerpt. The results below are this worked example's numbers, consistent with the b7 baselines and targets, and they are mixed: the primary metric improved and missed, the secondary barely moved, the guardrail held, and the decision rule fired. Read section 5 last and notice that nothing in it was decided today.

Read section 1 and find the shape. Staged rollout, not an A/B test. That admission constrains every claim in section 4, and it is in the first paragraph rather than a footnote.

Read section 2 before section 3. Cover the results with your hand if you have to. The prediction was 30%, the minimum worth shipping was 20%, and the decision rule is right there under them.

Now read section 3. 17.0% at week 3. The rule fired. Notice the readout says that in four words rather than in a paragraph about momentum.

Read the confound in section 4 twice. It is a genuinely good argument that the metric measured the wrong behaviour, and the readout refuses to use it as an excuse. That refusal is the hardest sentence in the document.

Check section 5 against section 2. Every action is the pre-committed rule being executed. If you can find one that was invented after the number arrived, the readout has failed.

READOUT  ·  Cadence post-meeting summaries  ·  staged rollout, weeks 1 to 4
Owner: PM        Status: decision rule fired        Shape: staged rollout with a guardrail

1 · WHAT WE SHIPPED
   Post-meeting summaries, on by default, for paid team workspaces of 5 to 50 seats.
   Staged: 20% of eligible workspaces in week 1, 60% in week 2, held at 60% to week 4.
   This was a staged rollout, not an A/B test. Anything below is observed, not causal.
   Live in-meeting notes did not ship. That trade-off was recorded in b6 and still holds.

2 · WHAT WE PREDICTED   (written 2 weeks before launch, unedited since)
   Primary    week-1 transcript revisit rate        11%    ->  30% by week 4
   Secondary  self-reported in-meeting note-taking  68%    ->  58% by week 6
   Guardrail  accuracy complaints                   1.4%   ->  stays at or below 1.4%
   Cost       roughly 0.02 per meeting
   Minimum effect worth shipping: 20% revisit rate. Below that, this is not worth a quarter.
   Decision rule: if revisit rate has not reached 20% by the end of week 3, we stop the
   phase-2 digest build and spend two weeks on the six highest-revisit accounts instead.

3 · WHAT HAPPENED
   Primary    revisit rate     17.0% end of week 3, 19.2% at week 4     MISSED  (target 30%)
   Secondary  note-taking      66% (n=142), was 68% (n=138)             NO MOVE (inside noise)
   Guardrail  complaints       1.3% of meetings                         HELD
   Cost       0.026 per meeting, about 30% over estimate                IMMATERIAL
   The decision rule fired at the end of week 3. 17.0% is not 20%. The week-6 note-taking
   read is not in yet; on the week-4 pulse it has not moved.

4 · WHAT WE NOW BELIEVE
   Summaries help, and they help less than we predicted. 8 points of revisit rate in four
   weeks is real, and it is not the 19 points we wrote down.
   Confound we can name: 71% of workspaces that opened a summary did not open the transcript
   in the same week. Our primary metric counts transcript opens, so it may be measuring the
   behaviour the feature was built to remove.
   We are not using that to call this a win. It was our metric, we chose it in b7, and we do
   not get to reinterpret it after seeing the number. It changes what we instrument next,
   not what we conclude this quarter.
   Confounds we cannot rule out: week 2 contained a public holiday in two large markets, and
   the in-product announcement banner may carry some of the lift on its own.
   What we did not expect: 9 of the 214 ticket-filing accounts asked for the summary as an
   email rather than in the app. Not a theme yet. Two sources short of one.

5 · WHAT WE ARE DOING NEXT
   a. Summaries stay on at 60% and go to 100% next week. The feature is not the problem.
   b. Phase 2, the digest email, is not funded this quarter. This is the rule firing, not a
      judgement made in this meeting.
   c. Two weeks on the six highest-revisit accounts: what did they do that the rest did not.
   d. Instrument a summary-open event before any new target is set. Next quarter's primary
      metric gets chosen from instrumented behaviour, not from this argument.
Job around the readoutHand it over?The line, and why it sits there
Draft the five-section structure from your rough notesyes, immediatelystructure is production work; the numbers stay yours and you paste them in
Argue the strongest opposite interpretation of your own datayes, and this is the highest-value use on the pageyou asked for the case against your conclusion, then you judged it - the judging is the job
List confounds you did not think ofyes, as a checklist generatorit is not an authority; you delete the ones that do not apply and check the ones that might
Rewrite the readout shorter for an exec audienceyessame substance, fewer words. If the substance changed, that is a rewrite failure, not a style choice
Decide what the result means for the roadmapneverthat is the decision, and the only thing allowed to make it is the rule you wrote in section 2
Pick the threshold now that the data is inneverthis is negotiating with yourself, with a fluent tool attached to make it sound like analysis
Real world

The confound in section 4 is the most tempting sentence a PM will write this year. It is true, it is well-argued, and it would let the launch be a success: if people read the summary instead of the transcript, then a transcript-open metric was always going to understate the feature. Every instinct says lead with it.

Leading with it costs you the only thing that makes readouts work. The metric was chosen in b7, in public, before anyone knew which way the number would go, and a metric you are allowed to reinterpret after seeing the result is not a metric - it is a mood with a baseline. So the confound goes where it belongs: as a named limitation, as the reason a new event gets instrumented, and as the reason next quarter's target is not set today. The result stays mixed. The credibility stays intact. In eighteen months, that credibility is the only reason anyone believes your next prediction.

Homework

Try it yourself - this week ◐ 25-35 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

Experiment practice - the falsifiable hypothesis, the decision rule, the stopping rule set in advanceParts 1 and 2 · the two-answer test and the four-part rule
The honest readout - prediction before result, the readout that can end in stopPart 3 and the worked artifact · Cadence's mixed result in full
Test shapes beyond the A/B test - staged rollout, painted door, conciergePart 2 · with what each shape cannot tell you
Minimum detectable effect and the minimum effect worth shippingDecision-shaped only · power and variance reduction are learn-experimentation
Metric definition and instrumentationUsed here, taught in b7 · driver-tree depth is learn-metric-decomposition
Rollout mechanics, release management, status reportingOut of scope by design - learn-ai-project-management
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · You write down what you would do if the test wins and what you would do if it loses, and the two answers are identical. What now?

The value of an experiment is exactly the difference between the two branches. When there is no difference, there is no value, and the instrumentation cost buys a chart for a launch email. Cancel it and say why in one line.

2 · The revisit rate hit 17.0% at the end of week 3 against a rule that said 20%. The team argues that week 2 had a holiday and asks to extend to week 5 before deciding. What is happening?

The holiday is genuinely a confound and belongs in section 4. It is not a reason to move a threshold, because every threshold moved after the fact moves to where the data already is. That is narration, not analysis.

3 · Which of these is the one job on this page that must never be handed to an AI tool?

Structure, counter-arguments and confound checklists are production work with a cheap verification step. What the result means for the roadmap is the decision, it belongs to the rule you wrote before the data, and it is the one thing on the page with your name on it.

PM session 8 cheat sheet · pin this

The two-answer testSame action whether it wins or loses? Do not run it. You are buying cover, not information.
A hypothesis can loseSegment, number, baseline, horizon, guardrail. If no result could contradict it, it is a slogan.
Decision rule, four partsMetric · threshold · date · the action on each side, including the side where you stop.
Before the dataStopping rule, minimum effect worth shipping, decision rule. Afterwards you negotiate with yourself.
Four shapesA/B · staged rollout with a guardrail · painted door · concierge on ten accounts. Name what yours cannot tell you.
Readout orderShipped · predicted · happened · now believe · doing next. Prediction before result, always.
The Cadence result19.2% against a 30% target, note-taking flat at 66%, guardrail held at 1.3%, rule fired.
Never delegateWhat the result means for the roadmap. Next: b9, the roadmap that has to survive three people.