Why the readout is the real skill
b7 set the numbers: week-1 transcript revisit rate as the primary metric, 11% today against a 30% target, self-reported in-meeting note-taking as the secondary at 68%, and accuracy complaints as the guardrail at or below 1.4% of meetings. The events are specified. Now something ships, and the organisation has to find out what happened without flattering itself.
Almost every team can design a test. Very few write a readout that is allowed to end in "we are stopping". That document is the one nobody is rewarded for, it is the one that saves the next quarter, and it is most of this session. The statistics side of experiments - power, variance reduction, sequential testing - is learn-experimentation, deliberately not re-taught here.
An experiment is for deciding, not for proving 7 min live
Start with the cheapest test in this whole course, and it is not a test at all. Write down what you would do if the result came back positive, and what you would do if it came back negative. If the two answers are the same, do not run the experiment. You already decided, and what you are about to buy is not information, it is cover.
LiveWriting a hypothesis that is allowed to lose4 min▶
A hypothesis is falsifiable or it is a slogan. Compare the two versions of the same belief about Cadence:
- Cannot lose: "Post-meeting summaries will improve the experience for team admins and drive engagement." Every observed outcome is consistent with this sentence. It cannot be wrong, so it cannot be informative, and the readout that follows it will be a paragraph of adjectives.
- Can lose: "If team admins on paid workspaces of 5 to 50 seats get a summary within 90 seconds of a call ending, week-1 transcript revisit rate rises from 11% to 30% within four weeks, while accuracy complaints stay at or below 1.4% of meetings." Named segment, named number, named baseline, named horizon, named guardrail. It can miss, and this one did.
Then attach the decision rule in the same breath, in the same document, with a date on it. Not "we will review the results" - the actual if-then, including the branch where you stop. If you cannot write the stopping branch, you are not designing a test; you are scheduling a celebration.
LiveThe decision rule, and where teams put it3 min▶
A decision rule has four parts, and a rule missing any one of them is a wish: the metric, the threshold, the date, and the action on each side. Cadence's rule in the figure above has all four, which is why it was able to fire without a meeting.
Put it somewhere a stakeholder will see it before the result exists. In the spec, in the launch note, in the channel where the launch is announced. A rule that lives only in your head is not a rule; it is a memory that will be revised by whatever the number turns out to be. The only real protection is that other people read the threshold while it was still uncomfortable.
And write the action on the losing side concretely enough that it can be executed by someone who is disappointed. "We stop the phase-2 digest build and spend two weeks on the six highest-revisit accounts" survives a bad Monday. "We will reconsider our approach" does not survive lunch.
Self-studyThree tests that should never have been run3 min read▶
- The already-decided test. The build is committed, the launch date is public, and the test is running to produce a chart for the launch email. Cost: three weeks of instrumentation and the credibility of every future test.
- The unpowered test on a small segment. A change tested on 40 accounts where only a very large effect would be visible, then reported as "no significant difference" and used to kill the idea. You did not learn it does not work. You learned nothing, expensively, and you learned it in a way that sounds like knowledge.
- The test nobody can act on. The result arrives four weeks after the planning cycle closed. Whatever it says, the roadmap is set. Run the test that lands before the decision, or move the decision, or do not run it.
All three failures happen before any data exists, which is the pattern of this session: the expensive mistakes in experimentation are design mistakes and readout mistakes, not analysis mistakes.
Three things you decide before a single event fires 6 min live
The order is the whole lesson. Set the stopping rule, the minimum effect worth shipping and the decision rule while the data does not exist yet, because once it exists you are no longer designing a test. You are negotiating with yourself, and you always win that negotiation.
LiveFour shapes, and three of them are not an A/B test4 min▶
A team of Cadence's size, with a segment of paid workspaces of 5 to 50 seats, often cannot run a well-powered A/B test on a metric like week-1 revisit rate inside a useful window. That is not a reason to skip the discipline. It is a reason to pick a shape that fits, and to be honest in the readout about what the shape can and cannot tell you.
| Shape | What it is | Use it when | What it cannot tell you |
|---|---|---|---|
| A/B test | randomised split, one change, a measured difference | you have the traffic, the change is isolated, and the decision can wait for the window | anything about accounts too rare to land in both arms |
| Staged rollout with a guardrail | 20% then 60% then all, with a pre-agreed number that pauses it | you are fairly confident, the risk is a regression rather than a dud, and speed matters | a clean causal estimate: the world moved while you rolled out |
| Painted door | the entry point exists, the feature does not; you measure intent to enter | the build is expensive and you are unsure anyone wants the door at all | whether they would keep using it once the room behind the door is real |
| Concierge test | ten accounts, the outcome delivered by hand, no product built | you need to know whether the outcome is valued before you automate it | anything about scale, cost, or self-serve discoverability |
Cadence shipped summaries as a staged rollout with a guardrail, and the readout below says so in its first paragraph. That single admission is what keeps the rest of the document honest, because it tells the reader which claims are causal and which are merely observed.
Self-studyThe concierge test nobody runs, and should3 min read▶
The CEO's agent ask has no evidence behind it. The instinct is to argue about it or to build a prototype, and both are expensive. The concierge shape is neither: pick ten workspaces, and for two weeks have a human do by hand whatever the agent would have done - chase the open action items, send the follow-up, update the tracker. No product. A shared doc and somebody's afternoon.
Two weeks later you know something nobody in the argument knew: whether admins let it happen without watching, what they corrected, and what they asked for that was not in the pitch. That is evidence, it cost no engineering, and it is exactly the discovery task b9 converts the CEO's ask into. Ten accounts and a human beats a quarter of building and a debate.
The five questions a readout answers 5 min live
A readout is not a summary of what happened. It is the document that converts a result into a decision, in front of witnesses, and its structure is the thing that makes honesty the path of least resistance. Answer these five in this order and the awkward parts have nowhere to hide.
LiveWhy prediction comes before result3 min▶
Section 2 exists to make section 4 hard to fake. If the reader sees the prediction first, "8 points of lift in four weeks" cannot be presented as a triumph when the number written down was 19 points. Put the result first and the prediction becomes optional context, then it becomes a footnote, then it disappears, and the readout becomes a highlight reel with a methodology section.
The same logic makes section 5 a consequence rather than a proposal. If the decision rule is already in section 2, then the action in section 5 is arithmetic that anyone in the room can check. If it is not, section 5 is a new argument you are making while everybody is emotional about the number, which is the worst possible time to make one.
Self-studyThe organisational problem behind bad readouts3 min read▶
People do not write dishonest readouts because they are dishonest. They write them because the last five honest ones were met with silence and the last five launch announcements were met with congratulations. The incentive is doing exactly what incentives do.
You cannot fix that alone, but you can do three things that measurably help. Write the decision rule publicly, so a mixed result reads as the system working rather than as your failure. Publish a readout for something that worked and include what you got wrong in it, so the format is not associated only with bad news. And when somebody else publishes a readout that ends in stop, say so out loud in the channel where the launch would have been celebrated. Three quarters of that and the format stops being career risk.
The actual readout for Cadence's summaries launch ★ mixed result, on purpose
This is the whole document, not an excerpt. The results below are this worked example's numbers, consistent with the b7 baselines and targets, and they are mixed: the primary metric improved and missed, the secondary barely moved, the guardrail held, and the decision rule fired. Read section 5 last and notice that nothing in it was decided today.
Read section 1 and find the shape. Staged rollout, not an A/B test. That admission constrains every claim in section 4, and it is in the first paragraph rather than a footnote.
Read section 2 before section 3. Cover the results with your hand if you have to. The prediction was 30%, the minimum worth shipping was 20%, and the decision rule is right there under them.
Now read section 3. 17.0% at week 3. The rule fired. Notice the readout says that in four words rather than in a paragraph about momentum.
Read the confound in section 4 twice. It is a genuinely good argument that the metric measured the wrong behaviour, and the readout refuses to use it as an excuse. That refusal is the hardest sentence in the document.
Check section 5 against section 2. Every action is the pre-committed rule being executed. If you can find one that was invented after the number arrived, the readout has failed.
READOUT · Cadence post-meeting summaries · staged rollout, weeks 1 to 4
Owner: PM Status: decision rule fired Shape: staged rollout with a guardrail
1 · WHAT WE SHIPPED
Post-meeting summaries, on by default, for paid team workspaces of 5 to 50 seats.
Staged: 20% of eligible workspaces in week 1, 60% in week 2, held at 60% to week 4.
This was a staged rollout, not an A/B test. Anything below is observed, not causal.
Live in-meeting notes did not ship. That trade-off was recorded in b6 and still holds.
2 · WHAT WE PREDICTED (written 2 weeks before launch, unedited since)
Primary week-1 transcript revisit rate 11% -> 30% by week 4
Secondary self-reported in-meeting note-taking 68% -> 58% by week 6
Guardrail accuracy complaints 1.4% -> stays at or below 1.4%
Cost roughly 0.02 per meeting
Minimum effect worth shipping: 20% revisit rate. Below that, this is not worth a quarter.
Decision rule: if revisit rate has not reached 20% by the end of week 3, we stop the
phase-2 digest build and spend two weeks on the six highest-revisit accounts instead.
3 · WHAT HAPPENED
Primary revisit rate 17.0% end of week 3, 19.2% at week 4 MISSED (target 30%)
Secondary note-taking 66% (n=142), was 68% (n=138) NO MOVE (inside noise)
Guardrail complaints 1.3% of meetings HELD
Cost 0.026 per meeting, about 30% over estimate IMMATERIAL
The decision rule fired at the end of week 3. 17.0% is not 20%. The week-6 note-taking
read is not in yet; on the week-4 pulse it has not moved.
4 · WHAT WE NOW BELIEVE
Summaries help, and they help less than we predicted. 8 points of revisit rate in four
weeks is real, and it is not the 19 points we wrote down.
Confound we can name: 71% of workspaces that opened a summary did not open the transcript
in the same week. Our primary metric counts transcript opens, so it may be measuring the
behaviour the feature was built to remove.
We are not using that to call this a win. It was our metric, we chose it in b7, and we do
not get to reinterpret it after seeing the number. It changes what we instrument next,
not what we conclude this quarter.
Confounds we cannot rule out: week 2 contained a public holiday in two large markets, and
the in-product announcement banner may carry some of the lift on its own.
What we did not expect: 9 of the 214 ticket-filing accounts asked for the summary as an
email rather than in the app. Not a theme yet. Two sources short of one.
5 · WHAT WE ARE DOING NEXT
a. Summaries stay on at 60% and go to 100% next week. The feature is not the problem.
b. Phase 2, the digest email, is not funded this quarter. This is the rule firing, not a
judgement made in this meeting.
c. Two weeks on the six highest-revisit accounts: what did they do that the rest did not.
d. Instrument a summary-open event before any new target is set. Next quarter's primary
metric gets chosen from instrumented behaviour, not from this argument.
| Job around the readout | Hand it over? | The line, and why it sits there |
|---|---|---|
| Draft the five-section structure from your rough notes | yes, immediately | structure is production work; the numbers stay yours and you paste them in |
| Argue the strongest opposite interpretation of your own data | yes, and this is the highest-value use on the page | you asked for the case against your conclusion, then you judged it - the judging is the job |
| List confounds you did not think of | yes, as a checklist generator | it is not an authority; you delete the ones that do not apply and check the ones that might |
| Rewrite the readout shorter for an exec audience | yes | same substance, fewer words. If the substance changed, that is a rewrite failure, not a style choice |
| Decide what the result means for the roadmap | never | that is the decision, and the only thing allowed to make it is the rule you wrote in section 2 |
| Pick the threshold now that the data is in | never | this is negotiating with yourself, with a fluent tool attached to make it sound like analysis |
The confound in section 4 is the most tempting sentence a PM will write this year. It is true, it is well-argued, and it would let the launch be a success: if people read the summary instead of the transcript, then a transcript-open metric was always going to understate the feature. Every instinct says lead with it.
Leading with it costs you the only thing that makes readouts work. The metric was chosen in b7, in public, before anyone knew which way the number would go, and a metric you are allowed to reinterpret after seeing the result is not a metric - it is a mood with a baseline. So the confound goes where it belongs: as a named limitation, as the reason a new event gets instrumented, and as the reason next quarter's target is not set today. The result stays mixed. The credibility stays intact. In eighteen months, that credibility is the only reason anyone believes your next prediction.
Try it yourself - this week ◐ 25-35 min total
- Run the two-answer test on the next thing you were going to measure: what you do if it wins, what you do if it loses. If the answers match, cancel the measurement and say why in one line.
- Write the decision rule for your carried backlog item: metric, threshold, date, and the action on each side. Post it where a stakeholder will read it while the threshold is still uncomfortable.
- Pick the shape honestly. If a well-powered A/B test does not fit your segment size, choose staged rollout, painted door or concierge, and write the sentence about what that shape cannot tell you.
- Take the last launch you were part of and write its readout retrospectively in the five sections. The gap between what was predicted and what got reported is your team's actual honesty budget.
- Draft one readout structure with AI, then ask it for the strongest argument against your own conclusion. Keep whatever survives your judgement and delete the rest yourself. Bring the roadmap consequence to b9.
Sources covered
Full source map in materials/official-course-map.md. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · You write down what you would do if the test wins and what you would do if it loses, and the two answers are identical. What now?
The value of an experiment is exactly the difference between the two branches. When there is no difference, there is no value, and the instrumentation cost buys a chart for a launch email. Cancel it and say why in one line.
2 · The revisit rate hit 17.0% at the end of week 3 against a rule that said 20%. The team argues that week 2 had a holiday and asks to extend to week 5 before deciding. What is happening?
The holiday is genuinely a confound and belongs in section 4. It is not a reason to move a threshold, because every threshold moved after the fact moves to where the data already is. That is narration, not analysis.
3 · Which of these is the one job on this page that must never be handed to an AI tool?
Structure, counter-arguments and confound checklists are production work with a cheap verification step. What the result means for the roadmap is the decision, it belongs to the rule you wrote before the data, and it is the one thing on the page with your name on it.