learn-ai-pm-with-phoebe / PM session 6 of 10
Learn AI Product Management with Phoebe · PM track · Session 6 of 10

Prioritisation and trade-offs

This is the session where the decision actually gets made. We score Cadence's three asks under two standard framings, and the top does not move but the margin collapses from 31x to 1.4x and second place flips to the ask with no evidence behind it. A ranking that changes with the lens was never an answer. So we do the other work instead: mark where every input came from, notice that confidence is both the riskiest input and the one people fake most, then make the call on the page, write what it cost, and record the branches we cut so the argument does not restart next quarter.

🟠 PM track PMs · founders · product leads The session where you choose 45 min
0-3 · Welcome 3-15 · Lenses, and the inputs they need 15-40 · Score twice, then choose 40-45 · Q&A
Part 0

Frameworks do not decide anything

By now the three asks are comparable. b2 gave each one an evidenced problem, b3 verified the research behind them, b4 named the job and wrote down the branches we cut, and b5 turned the leading candidate into something signable. Everything since b1 has been about making the asks sit in the same units. This session spends that work: two scoring lenses, one decision, and the sentence that says what the decision cost.

The reason prioritisation goes wrong is not that teams pick the wrong framework. It is that a framework returns a number, a number looks like an answer, and nobody asks where the inputs came from. AI makes this worse in a very specific way: ask it to score a backlog and it will fill in reach, impact, confidence and effort for all three items, in seconds, with two decimal places and no basis whatsoever. The output is arithmetic on guesses, and arithmetic is extremely convincing.

Live - presented in session Self-study - read after class ◆ Worked artifacts - real material to judge Sources covered
★ What you walk out with today The three asks scored under two lenses with a provenance column on every input, the decision made and written down, the trade-off in one sentence you could say out loud to the person who loses, and a cut record with re-open conditions.
Part 1 · covers prioritisation frameworks: RICE, WSJF, cost of delay

A framework is a lens, not an answer 7 min live

Two lenses, same three asks, same evidence, no cheating: RICE, which divides expected value by effort, and WSJF, which divides cost of delay by how long the job takes. Both are respectable. Both are used by serious teams. And they disagree, which is the useful part, because the disagreement tells you exactly what each lens is willing to ignore.

Same three asks. Same evidence. Two standard framings. Lens A · RICE = R x I x C ÷ E 1 · Post-meeting summaries 53.3 2 · Live in-meeting notes 9.3 3 · The agent 1.7 Effort sits in the denominator, so the 3-week estimate does most of the ranking. Confidence is an input, so weak evidence costs you points. Lens B · WSJF = (BV + TC + OE) ÷ weeks 1 · Post-meeting summaries 6.3 2 · The agent 4.6 3 · Live in-meeting notes 4.4 There is no slot for confidence anywhere in this formula, so an ask with no evidence at all scores on asserted value and urgency alone. The winner held. The margin did not: 31x clear under RICE, 1.4x under WSJF. And second place flipped to the only ask with no evidence behind it, because one lens has nowhere to record that.
🔍 Click to zoom - two respectable lenses, two different stories about the same three asks
LiveWhat each lens is willing to ignore4 min

Read the two formulas as statements of belief, because that is what they are:

  • RICE believes cheap reach wins. Effort is a divisor, so a 3-week job with broad reach beats an 11-week job almost regardless of how valuable the 11-week job is. It also has an explicit confidence term, which means it is the only one of the two that can express "we are not sure". That single term is why the agent finishes 31x behind rather than second.
  • WSJF believes urgency wins. Cost of delay rewards anything framed as time-critical or strategically enabling. It has no confidence input at all. So an ask with zero evidence and a strong narrative scores exactly as well as an ask with 214 tickets behind it, provided somebody says the words "market window".
  • Which is why lens choice is a decision, not a preliminary. Whoever picks the framework has already narrowed the answer. If someone arrives at a review with a ranking, the first question is not "what did it score" but "which lens, and who chose it".
The rule that comes out of this page Score under two lenses, always. If the ranking is stable, the frameworks have told you something. If it moves, they have told you nothing except which input is carrying the whole result, and that input is now the thing to go and check.
Self-studyWhy the score is not the artifact3 min read

Three things a score does well, and one it cannot do at all.

  • It forces the inputs into the open. You cannot score an ask without stating how many people it reaches and how long it takes, and half the value of the exercise is watching somebody discover they cannot state either.
  • It makes disagreement specific. "I think this is more important" becomes "I think your confidence of 0.5 should be 0.8, and here is why", which is a conversation that can actually finish.
  • It kills the bottom of the list cheaply. Items that lose under every lens are genuinely dead, and that is worth knowing early.
  • It cannot choose. A number ranks; a person commits. The gap between the two is the entire second half of this session, and no amount of scoring closes it.

The leader-track version of this is a4, prioritisation without theatre, which is about spotting a ranking built on guesses in somebody else's deck rather than building one yourself.

Part 2 · the inputs frameworks quietly require

Every score needs a provenance column 6 min live

Both formulas above quietly demand four things: reach, confidence, effort, and some notion of what waiting costs. None of the four arrives with the ask. All four have to be produced by somebody, and if you do not produce them, a drafting tool will, in a tone of complete assurance. The fix is small and it changes everything: one extra column that says whether each number was measured, estimated by someone who would know, or invented.

Four inputs. No framework tells you where any of them came from. Reach How many, in which segment, over what window. Needs a countable denominator. Confidence How sure you are the impact is real. Carries all the risk, and gets faked more than any. Effort Weeks, from the people who would build it. An estimate, always. Never a measurement. Cost of delay What a quarter of waiting costs. Least checkable of the four, so the most argued. Leave any of the four blank and it still gets filled in, to two decimal places, by something with no stake in the outcome. Three classes, and only three. Every number on the board gets one. Measured a system counted it. 214 tickets. 11% revisit rate. 0.02 per meeting. Estimated by a named person who would know. "3 engineer-weeks, from the tech lead who would build it." Invented nobody knows. Keep the number if you must, write the word, and stop reading that row as a score. An ask whose inputs are all invented has not been deprioritised by the maths. It has been excused by it. The provenance column is what stops a guess from graduating into a plan just because it went through a formula.
🔍 Click to zoom - four required inputs, three honest labels
Input, as used on this pageValueClassWhat would move it up a class
Reach · summaries100 of 100 recorded meetingsMeasurednothing needed: every recorded meeting on a paid workspace gets one, and analytics counts recorded meetings already
Reach · live notes68 of 100 meetingsMeasured, self-reportedthe 68% comes from a survey answer, not observed behaviour; instrument whether notes are taken elsewhere during a call
Reach · the agent30 of 100 meetingsInventedone interview in which an admin says what an agent would do for them, and how often
Impact · summaries2 of 3Estimated by someone who would knowthe support lead scored it against 214 tickets and 6 of 8 interviews; a ticket-to-account rollup would tighten it
Impact · live notes3 of 3Inventedpull the loss reasons on the 3 lost deals and check whether live notes was actually named; until then this is an assertion by an interested party
Confidence · summaries0.8Estimated by someone who would knowtwo independent signals agreeing (tickets and interviews) is what buys the 0.8; a third would not add much
Confidence · live notes0.5Estimated by someone who would knowone unverified signal, so half. Verify the deals and this moves in either direction, which is the point
Confidence · the agent0.25Estimated against nothingthe number a drafting tool writes here is 0.8, and the number the room says out loud is 0.8, and neither has anything under it
Effort · all three3 / 11 / 13 weeksEstimated by someone who would knowall three are engineering estimates and labelled as estimates on this page; the 11 includes a streaming pipeline that does not exist
Run cost · summaries0.02 per meetingMeasurednothing needed, and worth noticing that it is the only cost number on the board with a real basis
LiveConfidence is where the risk lives, and where the lying happens4 min

Of the four inputs, confidence is the strange one. Reach and effort are wrong by a factor of two at worst. Confidence is a multiplier on your entire belief that the thing works at all, and unlike the others it has no natural units, which makes it the easiest number in the whole exercise to move without anyone noticing.

  • It absorbs every unresolved question. Did the research cover the right segment? Was the sample friendly? Is the mechanism plausible? All of it lands in one decimal, and then the decimal gets multiplied by a big reach number and disappears.
  • It moves to fit the answer people want. Nobody writes 0.25 next to the CEO's request. Watch what happens to a confidence value between the draft and the review, and who was in the room when it changed.
  • A drafting tool always writes it high. Asked to score a backlog it produces confidence values in the 0.7 to 0.9 band across the board, because high confidence is what confident prose sounds like. It has not seen your research. It has not seen anything.
  • So state the basis, not the number. "0.8, two independent signals" and "0.8, because it feels likely" are the same cell and completely different claims. The provenance column is what makes them look different on the page.
LiveWhat to hand a drafting tool, and what to never let it touch3 min

There is real work for AI in this session, and it is all on one side of the line.

  • Genuinely useful: converting your inputs into both scorings without arithmetic slips, generating the sensitivity table (which single input, changed by how much, flips the ranking), writing the strongest possible case for the option you are about to cut, and drafting the note to sales once you have decided what it says.
  • Actively harmful: supplying any of the four inputs. Ask for a scored backlog and you get one, complete, plausible and untraceable, and the provenance column comes back full of numbers that belong in the invented row without saying so.
  • The tell: if every confidence value in a scored backlog sits between 0.7 and 0.9, nobody scored anything. Real confidence values are lumpy, because real evidence is lumpy.
The prompt shape that works Give it the inputs and the provenance labels you have already written, and ask for the arithmetic, the sensitivity analysis, and the argument against your own preference. Never ask it what the inputs are. The first is checking work; the second is outsourcing the judgement you are paid for.
Part 3 · the part no framework does

Then somebody has to choose 6 min live

The two lenses agree on the winner and disagree on everything else, so the honest reading is: summaries survived both framings, and the ordering below it is noise. That is as far as arithmetic goes. What follows is a choice, made by a person, on evidence, with a cost attached. Here it is, made on the page rather than left as an exercise.

The ranking narrowed it to a shortlist. Choosing is not a calculation. CHOSEN · Summaries 3 weeks. 0.02 a meeting. 214 tickets and 6 of 8 admins point at the same moment: after the meeting, not during. CUT · Live notes 11 weeks, and the streaming pipeline does not exist yet. Impact rests on 3 lost deals nobody has checked. CUT · The agent 13 weeks at the low end, and no evidence was offered. Not a no forever: a no until somebody names the job. The trade-off, written into the spec rather than the appendix: We are building post-meeting summaries and NOT live in-meeting notes this quarter. The pain the evidence describes happens after the meeting, and summaries take 3 weeks against 11. Sales loses the demo moment, and I am the one who says that to them. Then record the branches you cut, and the condition that would re-open each one. A cut with a reason and a re-open condition is a decision. A cut with neither is an argument that restarts in 12 weeks.
🔍 Click to zoom - the decision, the cost of it, and the record that stops the re-litigation
LiveWhy summaries, in the four sentences that actually did the work4 min

Neither score is the reason. These are:

  • The evidence locates the pain after the meeting. 214 tickets tagged "cannot find what was decided" and 6 of 8 interviewed admins who never reopen a transcript describe the same moment, and it is not the moment live notes serves. This is the argument. Everything else is supporting detail.
  • The two strongest signals are independent. Support tags and research interviews are different instruments pointed at different people, and they agree. That is what a confidence of 0.8 is actually made of.
  • It is the only option we can be wrong about cheaply. 3 weeks and 0.02 a meeting means being wrong costs a month. 11 weeks plus a pipeline that does not exist means being wrong costs the quarter and leaves infrastructure nobody asked for.
  • And it makes the live-notes question answerable sooner. Shipping summaries first buys eight weeks in which sales can pull the loss reasons on those three deals. If live notes is genuinely deal-blocking, that argument comes back stronger next cycle with a number in it.
The test for a real decision Can you name what you gave up, who loses by it, and what you will say to them? If the answer to any of the three is no, you have not decided. You have ranked, which is a different and much more comfortable activity.
Self-studySaying no to the CEO's agent without saying no to the CEO3 min read

The agent is the hardest of the three to cut, and not because it scores badly. It scores badly under both lenses. It is hard to cut because of who asked, which is a political problem wearing a prioritisation costume.

What does not work: presenting the score. A number that says the CEO's idea ranks third reads as a rebuttal, and the reply is that the framework is wrong, which is a debate you will lose because the framework is partly wrong, as this page just demonstrated.

What does work: convert the no into a missing input. "There is no version of this I can size, because nobody has told me which job it does for which admin. Give me one customer conversation where an agent is the obvious answer, and I will bring it back scored." That is not a refusal, it is a request, and the request is genuinely the blocker. If the conversation never happens, the ask dies of natural causes and nobody had to fight about it. If it does happen, you were wrong and now you have the input you were missing.

The stakeholder-facing version of this conversation, including the three questions the CEO will ask next, is b9.

Worked artifact · 12 min · everyone

The scored board and the decision record ◆ real material to judge

Both scorings on one board, with the provenance of the weakest input in the last column, because that column is the one that decides how much of the row you are allowed to believe. Read the rank-move column before the score columns.

Write the inputs down before you score anything. Reach with its denominator, impact on a stated scale, confidence as a decimal, effort in weeks from the people who would do the work. If a cell is empty, leave it empty and see who volunteers a number.

Label every cell: measured, estimated by someone who would know, or invented. Do this before the arithmetic, not after, or the labels will drift towards whatever justifies the ranking you already have.

Score under two lenses. Then compare the rankings, and more importantly compare the margins. A winner that holds under both lenses but goes from 31x clear to 1.4x clear is telling you the gap is a modelling artifact.

Find the single input carrying the result. Change one cell at a time and watch the order. Here it is effort: move the summaries estimate from 3 weeks to 9 and the RICE score falls from 53.3 to 17.8, still first. Move it to 30 and it loses. So the decision survives a 3x estimation error and not a 10x one, which is worth knowing.

Then stop scoring and write the record. The decision, the reason, what it cost, what you will say to the person who lost, the cut branches, and the condition that re-opens each one. This artifact outlives the spreadsheet by about a year.

AskRICE inputs
R · I · C ÷ E
RICEWSJF inputs
BV + TC + OE ÷ weeks
WSJFRank moveWeakest input, and where that number came from
Post-meeting summaries100 · 2 · 0.8 ÷ 353.313 + 3 + 3 ÷ 36.31 → 1Effort, 3 weeks: estimated by the tech lead who would build it. The 0.02 per meeting run cost is measured, which makes this the best-sourced row on the board
Live in-meeting notes68 · 3 · 0.5 ÷ 119.320 + 20 + 8 ÷ 114.42 → 3Impact, scored a generous 3 of 3: invented. It rests entirely on 3 lost deals that nobody has checked against the loss reasons, asserted by the team that wants the feature
The agent30 · 3 · 0.25 ÷ 131.720 + 20 + 20 ÷ 134.63 → 2Reach, 30 of 100: invented, and so are impact and all three cost-of-delay components. Five invented cells producing a second-place finish under lens B

Note what the WSJF column did to the bottom of the board. The ask with the most invented inputs finished second, ahead of the ask with one unverified signal, because the lens has no confidence term and rewards asserted urgency. Nothing dishonest happened. Somebody just chose a framework.

DECISION RECORD  ·  Cadence  ·  next quarter  ·  owner: PM  ·  reviewed with: eng lead, support lead, sales lead

WE ARE BUILDING    Post-meeting summaries, for team admins on paid workspaces of
                   5 to 50 seats. Signable spec: see b5, scored 100 of 100.

BECAUSE            The evidence locates the pain AFTER the meeting rather than during
                   it. 214 support tickets in 3 months tagged "cannot find what was
                   decided", the second most common tag; and 6 of 8 interviewed admins
                   never reopen a transcript at all. Two independent instruments,
                   same finding.

AT THE COST OF     Live in-meeting notes. This ships without the demo moment sales
                   asked for, and we are accepting that for at least one quarter.

WHAT I WILL SAY    "Your three lost deals are the strongest thing anybody brought me,
TO SALES           and I could not check a single one of them. So here is the deal:
                   pull the loss reasons on those three, and flag any deal this quarter
                   where live notes gets named, and I will re-score it next cycle with
                   your number instead of my guess. Meanwhile the thing 214 tickets are
                   about takes 3 weeks rather than 11, which means your argument gets
                   its hearing sooner this way than if I started live notes today."

CUT, AND WHY       Live in-meeting notes  ·  impact rests on 3 unverified deals, and it
                   needs a streaming pipeline that does not exist. 11 weeks.
                   The agent  ·  no evidence was offered, and no job has been named.
                   13 weeks at the low end.

RE-OPEN IF         The loss-reason pull names live notes in 3 or more deals; or an admin
                   conversation names a specific job an agent does; or the summaries
                   metric set in b7 does not move.

WHAT WOULD HAVE    A verified deal-loss number for live notes. Reach and impact for the
CHANGED THIS       agent that are not invented. A summaries effort estimate above
                   30 weeks.
Real world

Lens shopping is the most common form of prioritisation theatre, and it never looks like cheating. A team runs RICE, does not like the answer, and someone points out that RICE undervalues strategic bets. Which is true. So they run WSJF, and the strategic bet comes second instead of last, and the deck says "scored under two frameworks" as though that were rigour rather than the opposite. The countermeasure is not banning frameworks. It is fixing the lens and the provenance labels before anybody sees a ranking, and writing down which lens you chose and why in the same document as the score. A ranking is only evidence if the method was picked before the answer was known.

Practice

Your turn: three passes over the board ◆ 8 min

No tool needed. A pen and the table above.

LiveQ1 · Re-score the agent with honest confidence3 min

The agent's WSJF row is 20 + 20 + 20 ÷ 13 = 4.6, second place. Every one of those three 20s is invented, and WSJF has no confidence term to catch that. So add one by hand: multiply the cost of delay by the confidence you would actually defend, which for an ask with no evidence at all is 0.25.

60 x 0.25 = 15, over 13 weeks, gives 1.2. Now it is last under both lenses, by a wide margin, and it got there without anyone arguing about whether agents matter. That is the move: when a framework has no slot for how much you believe the inputs, add the slot rather than accepting the ranking. Then say so in the document, because you have just modified the method and that is a thing reviewers are entitled to know.

LiveQ2 · Find the input doing all the work3 min

Take the summaries row and change one input at a time until the ranking breaks. Then do the same for live notes.

You will find two things. First, the summaries decision is robust to reach and impact being wrong by a lot, and fragile to the effort estimate being wrong by a factor of ten, which is unlikely for a 3-week job and very likely for an 11-week one. Second, live notes flips ahead of summaries under RICE if its confidence goes from 0.5 to 0.9 and its effort comes in at 4 weeks rather than 11. So the whole live-notes case reduces to two claims: the deals are real, and the pipeline is easier than we think. Both are checkable, neither has been checked, and now they are written down as the two things that would change the decision.

This is the sensitivity pass, and it is the one part of the session worth handing to a tool: give it your inputs and ask which single cell, changed by how much, flips the order. It is arithmetic on numbers you supplied, which is exactly the safe half of the line.

Self-studyQ3 · Audit a ranking you did not build4 min read

Find the most recent prioritised backlog in your company that you did not personally score. Add one column: measured, estimated by someone who would know, or invented. Do not change any numbers and do not tell anyone yet.

Two patterns show up almost every time. The confidence column is uniformly high, which means nobody scored confidence, they scored enthusiasm. And the items at the top with the smallest effort estimates were estimated by the person who wanted them built, not by the person who would build them. Neither finding requires a confrontation. Both can be raised as one question in the next planning meeting: "can we mark which of these numbers we measured?" It reads as tidiness, it is answerable, and the answer reorders the list.

Homework

Try it yourself - this week ◐ 25-35 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

Prioritisation frameworks: RICE, WSJF, cost of delay, as lenses rather than answersPart 1 + the worked board · same asks, two framings, the margin and the ordering both move
The inputs frameworks quietly require, and the provenance discipline that keeps them honestPart 2 · three classes only, with confidence called out as the one that carries the risk
Writing down the branch you cut, from opportunity-tree practicePart 3 + the decision record · every cut gets a reason and a re-open condition
Metric definition, baselines and targetsNamed here as the thing the re-open condition points at - the full treatment is b7
The six spec checksReferenced, not re-taught - b5 owns them, and the trade-off is check 6
Estimation practice, sequencing, and anything after the decisionOut of scope by design - learn-ai-project-management
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · The agent finishes third under RICE (1.7) and second under WSJF (4.6). What changed?

RICE multiplies by confidence, which is the only place either formula can record "we have nothing behind this". Remove that term and five invented cells produce a second-place finish. Neither framework is wrong; they disagree about what is allowed to be ignored, and that disagreement is the finding.

2 · You need a reach number for an ask and nobody has one. What do you do?

The problem is never that a number is missing, it is that a missing number gets replaced silently and then survives into a plan. Labelling it invented keeps the row on the board, keeps the gap visible, and turns the next step into a small research task rather than an argument about priorities.

3 · Summaries win under both lenses. Has the framework made the decision?

A number ranks; a person commits. The test for whether you decided rather than ranked is whether you can name what you gave up, who loses by it, and what you will say to them. Nothing in either formula produces any of those three.

PM session 6 cheat sheet · pin this

Lenses, not answersScore under two. Stable ranking = a finding. Moving ranking = one input is carrying everything.
The Cadence boardRICE: 53.3 / 9.3 / 1.7. WSJF: 6.3 / 4.6 / 4.4. Same winner, 31x margin down to 1.4x.
Four required inputsReach, confidence, effort, cost of delay. None arrives with the ask. All four get invented if you leave them blank.
Three provenance classesMeasured · estimated by someone who would know · invented. Only three, and every cell gets one.
Confidence is the riskNo natural units, absorbs every unresolved question, moves to fit the answer people want. State the basis, not the decimal.
The AI lineGive it your inputs, get arithmetic, sensitivity and the counter-argument. Never ask it what the inputs are.
Test for a real decisionName what you gave up, who loses, and what you will say to them. Three noes means you ranked, not decided.
The cut recordReason plus re-open condition, or the argument restarts in 12 weeks. Next: b7, the metric before the build.