learn-ai-pm-with-phoebe / Leader session 5 of 6
Learn AI Product Management with Phoebe · Leader track · Session 5 of 6

Metrics and the honest readout

The mechanics take fifteen minutes: one primary metric per decision, a leading number if you want to be able to act, a guardrail as the promise to everyone else, and a baseline you almost certainly do not have. The rest of this session is the part that actually determines whether any of it works, and it is not technical. It is whether your organisation rewards a readout that says the bet did not work. If it does not, your PMs will keep producing metrics, and none of them will ever tell you anything you did not want to hear.

🟡 Leader track Heads of product · founders · stakeholders No code to write 45 min
0-3 · Welcome 3-17 · One metric, leading or lagging 17-31 · Does your org reward a no? 31-45 · Discussion + the questions to ask
Part 0

Why this session exists

Cadence's week-1 transcript revisit rate is 11% of recorded meetings. Nobody has ever set a target for it. That is not a story about a badly run product; it is the normal state of most product metrics in most companies, including yours, and it is why so many quarters end with an argument about whether something worked rather than a decision about what to do next. A number with no baseline cannot be moved. A number with no target cannot be missed. And a number with no date attached will be read whenever the news is good.

Session a1 set the production and commitment split, a2 the evidence standard, a3 what a spec must contain before it is signable, and a4 how to read a ranking without being captured by it. This one covers the number attached to the decision, and then the harder question underneath: what happens in your organisation when that number comes back flat. As always, it ends with the questions to ask your PMs.

Live - discussed in session Self-study - read after class ◆ Discussion exercise - run it in your own review Sources covered
★ What you walk out with today One primary metric per decision as a rule you enforce, the leading-or-lagging test applied out loud, an audit of your own last four readouts, and a decision about how you will respond to the next honest no.
Part 1 · covers North Star and guardrail metric practice

One primary metric per decision 7 min live

This is the rule that saves the most arguments, and it is the one most often negotiated away in the meeting where it gets set. Three success metrics feels more rigorous than one. It is the opposite: it is the absence of a commitment, formatted as thoroughness. Measure as many things as you like, and instrument generously. But exactly one of them is the number the decision is judged on, and it is chosen before the build starts.

One decision, one primary metric. The alternative is not more rigour. One primary metric · revisit rate: 11% today, 30% in 90 days · if it moves, the bet was right · everything else is context or a guardrail · one number to be wrong about, agreed early · the readout more or less writes itself Three metrics of equal weight · revisit rate up, activation flat, NPS down · every outcome is a win for somebody · the readout becomes a negotiation · whichever moved gets promoted to primary · nobody was wrong, which is the problem More than one primary is the same as none, and the reason is not that measuring three things is wrong. Measure twenty. The primary is the one number you agreed in advance to be judged on, and you can be judged on only one thing at a time. If three rank equally, you will pick whichever moved, and the readout stops being evidence.
🔍 Click to zoom - why three co-equal metrics is a commitment nobody made
LiveThe baseline nobody has4 min

Cadence's revisit rate sits at 11%, with no target ever set. Look at what that single fact makes impossible:

  • You cannot tell success from noise. If it reads 13% after the build, was that the build, a seasonal effect, or the weekly variance nobody has ever measured? Without a history, 13% is a number rather than a result.
  • You cannot size the opportunity in advance. Is 30% ambitious or trivially achievable? Nobody in the room can say, which means the target will be set by whoever sounds most confident.
  • You cannot fail. This is the one that matters. A metric with no target cannot be missed, so the build cannot be wrong, so the quarter ends with a discussion about interpretation instead of a decision about what to do next.

The fix is not sophisticated. Somebody spends an afternoon pulling the number and its variance over the last few months, before anything gets built. The reason it does not happen is that this work is invisible, unglamorous, and always looks postponable, right up until the readout when it is the only thing anybody needs.

The two-part rule you can enforce from tomorrow No target without a baseline, and no baseline that arrives after the build. If the number does not exist before you start, you are not running a bet, you are running an activity, and it will be judged by whoever tells the best story about it afterwards. The PM track instruments this before the build in b7, deliberately in that order.
Self-studyThe guardrail as a promise to the rest of the company3 min read

A guardrail is not a secondary success metric. It is the number that says: in pursuing our thing, we promise not to damage yours. That makes it a political instrument as much as a measurement one, and it is the cheapest trust you will ever buy across teams.

On Cadence, the guardrail is transcript accuracy complaints, currently 1.4% of meetings, and the promise is that it does not rise. Support owns the consequences of that number, and support did not choose to run this experiment. Naming the guardrail in advance converts a vague worry into a specific commitment, and it means the summaries team can move quickly without the rest of the company having to trust them personally.

Three properties a real guardrail has: somebody other than the builder cares about it; it has a threshold rather than a direction, so "must not exceed 1.4%" and not "keep quality high"; and breaching it stops the work rather than triggering a discussion. Guardrails that only trigger a discussion are decoration, and everyone learns that within a quarter.

Part 2 · covers leading, lagging and guardrail metrics

Can it steer, or only judge? 7 min live

Most explanations of leading versus lagging are taxonomy and change nobody's behaviour. Here is the version that does. A metric that moves within a week can change what you build next week. A metric that moves at quarter end can only tell you afterwards whether you were right. Both are legitimate, they are not substitutes, and confusing them is how teams end up flying an entire quarter with no instruments and a verdict waiting at the end.

Same build, three metrics, three completely different uses week 1 week 2 week 4 quarter end Leading moves within a week, so it can still change what you build Lagging moves at quarter end, so it can only judge the build, never steer it Guardrail watched the whole time; it protects everyone else, it does not prove you right On Cadence: transcript revisit rate moves within a week, which makes it steerable. Seat retention moves in a quarter. The guardrail is accuracy complaints, currently 1.4% of meetings, and the rule is that it must not rise. Pick a leading primary if you want to be able to act, and expect the board to ask about the lagging one anyway. Both are true at once.
🔍 Click to zoom - the only distinction that changes behaviour is when the number arrives
TypeWhat it can doWhat it cannot doCadence example
Leading
moves within a week
steer a build mid-flight; kill a bet early and cheaply; give a team feedback while they can still act on itprove the business result, or survive as the only number a board wants to talk aboutweek-1 transcript revisit rate: 11% today, 30% target in 90 days
Lagging
moves in a quarter or more
judge whether the bet paid; anchor the board conversation; settle whether a strategy is workingtell you anything in time to change the build, or distinguish your work from four other changespaid seat retention on workspaces of 5 to 50 seats
Guardrail
watched continuously
stop a win that costs another team more than it gains; buy trust across the org in advancetell you whether the bet worked; it is a promise, not a resulttranscript accuracy complaints at 1.4% of meetings, must not rise
LiveWhat to do when the two disagree3 min

They will, and the moment is more common than the textbooks suggest. Revisit rate climbs from 11% to 24% and seat retention does not move at all. Now what?

  • The leading metric was your bet about causation, and it just failed. You believed people who revisit transcripts renew. They revisited and did not renew. That is a genuinely valuable finding about your business, and it is worth more than the feature was.
  • Do not quietly swap to the metric that moved. This is the most common dishonest readout in product work, it is almost never done cynically, and it destroys the value of every future readout in the same organisation.
  • Say both numbers, in that order, with the original target attached. "We hit the leading metric, we missed the lagging one, so our model of why this mattered was wrong" is a strong readout. It is also the sentence most orgs cannot say out loud, which is the subject of Part 3.

Note the boundary: how much movement counts as movement, and how confident you can be, is statistics rather than product judgement. Deciding the threshold in advance is your job. Power, variance and sequential testing are learn-experimentation, and breaking a metric into its drivers is learn-metric-decomposition.

Self-studyWhy AI made this worse, briefly2 min read

Two specific effects, both small-sounding and both real.

First, "improve engagement" is now free to write, formatted convincingly, in a success section that looks complete. A drafted spec will hand you a directional phrase where a metric should be, and a reviewer skimming for structure will approve it. That is the absence of a metric wearing the shape of one, and it is the failure the PM track's spec scorer catches as check 3.

Second, and more subtly, a fluent narrative can now be produced around any set of numbers in about a minute. Post-hoc justification used to cost effort, and that effort was a natural brake. It costs nothing now. So the guard has to move upstream, into the target and the date being fixed before the build, because a well-written explanation of a flat metric is no longer evidence of anything at all.

Part 3 · covers stopping rules and the honest readout

Does your org reward a no? 10 min live

Everything above is mechanics, and mechanics are the easy half. A team with perfect metric hygiene inside an organisation that punishes bad news will produce beautifully instrumented readouts that are all, somehow, positive. You will not be lied to. You will simply find that the framing is always generous, the metric that moved is always the one discussed, and losses arrive labelled as learnings with no numbers attached. That pattern is not a character problem in your PMs. It is a reporting line responding to incentives, and the incentives are yours.

Read the left column honestly. Most organisations have at least three of them. Tells that a no is not safe here · readouts appear when results are good, slip when not · the metric quietly changes between plan and readout · a loss is written up as learnings, with no number · nobody can name a bet that was stopped last year · the PM asks you what you want the answer to be Five moves, one quarter, no budget · ask for the stopping rule before the build starts · fix the readout date in advance, in writing · read it on that date even when it looks bad · say out loud which bets you expect to fail · promote someone who reported a clean loss The culture is set by what you do with the first honest no, and everybody hears about it within a day. Not by what is on your wall. Thank the PM, ask what they learned, hand them the next bet, and you have an honest readout culture. Ask what went wrong with them and you will not get another. The PM track's b8 writes exactly that readout, with a mixed result.
🔍 Click to zoom - the diagnosis, the five moves, and the one moment that decides it
LiveThe stopping rule, and why it has to come first4 min

A stopping rule is one sentence, written before the build, that says what result would make you abandon this. "If revisit rate is under 18% at four weeks, we stop and do not build phase two." It is the highest-leverage sentence in the whole metrics conversation, for a reason that has nothing to do with measurement.

  • Written first, it is a technical judgement. Nobody is invested, nobody has spent a quarter, and a reasonable threshold is easy to agree.
  • Written afterwards, it is a verdict on a person. By then the team has shipped, the number is known, and any threshold anybody proposes is obviously chosen to produce a particular answer. This is why the rule cannot be added later, no matter how good everyone's intentions are.
  • It converts stopping from a failure into compliance. A team that stops because the pre-agreed rule fired has followed the process correctly. That distinction is the entire difference between an org that can kill work and an org where everything ships.
Ask for it in the approval meeting, not the review "What result would make us stop?" costs you eight seconds in the meeting where you approve the work, and it is nearly free to answer at that point. In the review it is an accusation. Same question, completely different instrument, depending entirely on when you ask it.
LiveWhat a good honest readout actually looks like3 min

Short, numerate, and unembarrassed. Five lines, roughly:

  • What we believed: that admins do not revisit transcripts because transcripts are hard to read, and a summary would fix it.
  • What we agreed in advance: revisit rate from 11% to 30% in 90 days, stopping rule at 18% by week four, guardrail at accuracy complaints not rising above 1.4%.
  • What happened: revisit rate reached 24%, retention did not move, complaints stayed at 1.3%. Mixed, and stated in that order rather than reordered to lead with the win.
  • What we now believe: revisiting was a symptom, not a driver. The renewal conversation is somewhere else.
  • What we are doing next: not phase two. One week of interviews with the accounts that did renew.

Notice that this readout contains a loss and is not apologetic, and notice there is no paragraph explaining why the target was always ambitious. The confidence comes from the numbers having been fixed in advance. That is the whole mechanism: you cannot be accused of moving the goalposts if you wrote them down and left them there.

Self-studyThe learnings tell, and how to close it2 min read

Watch for the sentence "we learned a lot from this" arriving without a single number. It is almost always a loss being reported in a register that cannot be challenged, and it works precisely because nobody wants to be the person who objects to learning.

The close is polite and immediate: "Great, what specifically did we learn, and what will we do differently because of it?" A real learning is a changed belief with a consequence attached, and it takes one sentence. A cover story cannot answer the second half, because there is no consequence, because nothing was actually learned other than that the bet did not pay.

Ask it warmly and ask it every time, including on the wins. If you only ask it after bad results it becomes a punishment signal, and you have taught your team that the word learnings triggers an interrogation, which is the opposite of what you wanted.

Discussion exercise · 12 min · everyone

Audit your own last four readouts ◆ run this in your next review

The first two prompts belong in your next approval meeting and cost nothing. The third one is the real exercise, it is uncomfortable, and it is about your organisation rather than any individual. Run it on yourself before you run it on anybody else.

Everyone in the room finds their last four readouts from real work, in whatever form they exist. If they do not exist in any form, that is the first finding and the exercise is already worth the time.

Count how many said the bet did not work. Just the count. Zero out of four is the most common answer and it is not because the bets were good.

Check each one for a moved goalpost. Compare the metric named in the readout with the metric named in the original plan. Any difference is the finding, and it is usually not deliberate.

Run prompts 1 and 2 on the next thing you approve. Both live in the approval meeting, not the review. That timing is the whole trick.

Decide, now, what you will do with the next honest no and say it out loud in this room. A commitment made in advance is the only kind that survives the meeting where the news is bad.

LivePrompt 1 · the number, before the build4 min
Say this, word for word "Before this starts: what is the one number, what is it today, what will it be in ninety days, and on what date do we read it? Four answers, one line each. If you do not know what it is today, say so and tell me when you will."

What a good answer sounds like: four specific answers, and a PM who knows where the current value came from. "Revisit rate. It is 11% now, I pulled it Monday. Thirty percent at ninety days. We read it on the 14th of next month." An excellent answer adds the guardrail unprompted, and an honest one says "I do not have today's value yet, I will have it Thursday" rather than inventing a plausible baseline on the spot.

The failure mode to listen for: the answer is directional. "We will look at engagement" or "activation should improve" is the absence of a metric, formatted as one. Second tell: no current value, which means the target was set by feel and cannot be missed. Third and most important: the read date is "after launch" or "once it settles". An unfixed date is what allows a readout to appear only when the news is good, and it is the cheapest thing on this list to fix.

LivePrompt 2 · the stopping rule4 min
Say this, word for word "What result would make us stop this? Give me a number and a date, now, and write it into the spec. I am asking before we start precisely so that stopping is not a failure if it happens."

What a good answer sounds like: a threshold and a date, offered with slight discomfort. "If revisit rate is under 18% at four weeks, we stop and do not build phase two." The discomfort is correct and you should say so, because the PM has just made themselves accountable to a number in front of you. The second sentence of the prompt is the part that makes the answer possible, so do not drop it.

The failure mode to listen for: "we would look at the data and decide." That is not a stopping rule, it is a plan to never stop, and it means the work will continue until it is finished regardless of what happens. Also listen for a threshold set so low it cannot be breached, which is the same thing with better manners. If the rule could not plausibly fire, you do not have one.

Self-studyPrompt 3 · the readout audit, asked of your own team4 min
Say this, word for word "Pull our last four readouts. Which of them said the bet did not work? If the answer is none, I want your honest read on why, and I am asking about how I run this team rather than about your work."

What a good answer sounds like: somebody names one, with the number in it. Or, better and rarer, somebody says "none of them, and I think it is because the last person who reported a flat quarter got a lot of questions about their judgement." That answer is a gift and you should treat it as one out loud, immediately, because the room is watching how you take it.

The failure mode to listen for: "all four worked." Four out of four is not a strong team, it is a reporting system with a filter in it. Second tell: the metric in the readout is not the metric in the plan, and nobody noticed. Third: everyone looks at you before answering, which tells you the answer is being calibrated to your face rather than to the record. If that happens, the honest move is to say what you are going to change about your own behaviour, and then do it on the next flat result.

Real world

Nobody sets a target on 11%, and everybody has an 11%. Cadence's transcript revisit rate had been sitting at 11% of recorded meetings for as long as analytics had existed, on a chart nobody looked at, with no target and no owner. It was not hidden and it was not disputed. It was simply never anybody's number, which is the normal condition of most product metrics in most companies. The consequence surfaces at exactly one moment: the readout. Ship a summaries feature, watch revisit rate read 14%, and now the room has to decide whether that is a win, and it cannot, because nobody knows the weekly variance, the seasonal shape, or what the number did the last three times the product changed. The whole quarter of work resolves into a conversation about interpretation. The fix costs one analyst one afternoon, before the build, and the only reason it does not happen is that it looks postponable and nothing bad happens for three months.

Take this to your next review

The questions to ask your PMs ◐ 7 questions

The first four belong in the meeting where you approve work, where they are cheap and nearly free to answer. The last three belong in the review, and two of them are questions about you rather than about them. Ask those two warmly or do not ask them at all.

Homework

Try it yourself - this week ◐ 30-40 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

North Star and HEART-style metric practice: one primary metric per decision, guardrailsParts 1 and 2 · the primary, the baseline, the promise to other teams
Experiment practice: stopping rules decided in advance, and the honest readoutPart 3 + the discussion · the PM track's b8 writes one, with a mixed result
AI-assisted knowledge work: directional metrics and free post-hoc narrativePart 2 self-study · why the guard has to move upstream of the build
Instrumentation and event specs written before the buildNamed here, taught in the PM track's b7
Statistical depth: power, minimum detectable effect, variance, sequential testingOut of scope by design - learn-experimentation
Breaking a metric into its driversOut of scope by design - learn-metric-decomposition
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · A PM proposes three success metrics of equal weight for one decision. What is the actual problem?

Measure twenty things. The primary is the single number you agreed in advance to be judged on, and you can only be judged on one thing at a time. Three co-equal metrics is not extra rigour, it is the absence of a commitment formatted as thoroughness.

2 · The only metric attached to a build moves at quarter end. What have you actually got?

A metric that moves in a week can change what you build next week. A metric that moves in a quarter can only tell you afterwards whether you were right. Both are legitimate and they are not substitutes, so a team with only a lagging number is flying the quarter with no instruments.

3 · A PM's readout says the bet did not work, with the pre-agreed numbers attached. Which response actually changes your culture?

The culture is set by what happens to the first person who says it did not work, and everybody hears about it within a day. Handling it quietly is nearly as damaging as handling it badly, because the org learns that honest results are something to be absorbed rather than rewarded.

Leader session 5 cheat sheet · pin this

One primary per decisionMore than one is the same as none: you will pick whichever moved, and the readout stops being evidence.
Leading or laggingA week means it can steer the build. A quarter means it can only judge it. Not substitutes.
The baseline nobody hasCadence: revisit rate 11%, no target ever set. Normal, and it is why quarters end in interpretation.
The guardrailA promise to another team, with a threshold, that stops the work. Cadence: complaints at 1.4%, must not rise.
The stopping ruleWritten first it is technical. Written afterwards it is a verdict on a person. Ask in the approval meeting.
Fix the date and the audienceA read date set in advance is what stops readouts appearing only when the news is good.
The learnings tell"We learned a lot" with no number is a loss in an unchallengeable register. Ask what changes because of it.
The one moment that mattersWhat you do with the first honest no. The PM track's b8 writes one, mixed result. Next: a6, the org and the budget.