learn-ai-pm-with-phoebe / Leader session 4 of 6
Learn AI Product Management with Phoebe · Leader track · Session 4 of 6

Prioritisation without theatre

Rank three options one way and summaries win. Rank the same three another way and live notes win. Nothing about the options changed, which tells you the framework never decided anything. It made a decision legible, and the deciding happened when somebody chose the lens. This session is about the inputs those frameworks quietly require, the fact that a drafting tool will supply every one of them as a confident number with no basis, and what you should require instead of a spreadsheet with two decimal places.

🟡 Leader track Heads of product · founders · stakeholders No code to write 45 min
0-3 · Welcome 3-16 · Frameworks are lenses 16-30 · Provenance and confidence 30-45 · Discussion + the questions to ask
Part 0

Why this session exists

You have sat in the meeting. A PM presents a scored backlog, the scores run to two decimal places, the top item wins by a fraction, and everybody nods because the process was followed. Then somebody asks where the reach number came from and the room goes quiet. That silence is the subject of this session, and it got louder the moment drafting became free: a tool will now fill in reach, confidence, effort and cost of delay for you, instantly, in the same confident register as the numbers that came out of your warehouse.

Session a1 set the production and commitment split, a2 set the evidence standard, and a3 covered what a spec has to contain before anyone signs it. This session is the ranking meeting: how to read a score without being captured by it, what to demand about where each input came from, and what to accept as the record of the decision. It ends where every session in this track ends, with the questions to ask your PMs.

Live - discussed in session Self-study - read after class ◆ Discussion exercise - run it in your own review Sources covered
★ What you walk out with today A provenance rule you can impose on the next ranked backlog you are shown, one question that grades a score in three words, and the three lines you will accept instead of a spreadsheet.
Part 1 · covers prioritisation frameworks as lenses

The same options, ranked twice 11 min live

Cadence has three asks on the table: live in-meeting notes, which sales wants because every demo asks for it and three deals were lost; post-meeting summaries, which support wants because 214 tickets in three months are tagged "cannot find what was decided"; and an agent, which the CEO wants and for which no evidence has been offered. Below, one PM ranks those three twice in an afternoon. Both rankings are competent. Both are arithmetically correct. They disagree about what to build.

The same three Cadence asks, ranked twice, by the same PM on the same day Lens A: reach x impact x confidence / effort 1 · Post-meeting summaries 8.4 2 · Live in-meeting notes 6.1 3 · An agent 2.0 Effort sits in the denominator, so the known, cheap build wins. Nothing prices the 3 lost deals. Lens B: cost of delay / duration 1 · Live in-meeting notes 9.0 2 · Post-meeting summaries 7.5 3 · An agent 3.0 Cost of delay sits in the numerator, so the revenue story wins. Nothing prices the 214 tickets. Same three options, same evidence, two orders. The framework did not decide anything. It made a decision legible, which is worth doing and is not the same as being objective. Choosing the lens IS the decision, and a person makes it before any number is entered. The PM track makes exactly this call on Cadence in b6.
🔍 Click to zoom - two lenses, two orders, one decision that was made before the arithmetic
LiveWhat a framework is actually for4 min

None of this makes scoring useless. It makes scoring a different thing than most teams present it as, and the difference matters to you specifically, because you are the one being asked to approve the output.

  • A framework makes a decision legible. It forces every option into the same units, so the comparison can be argued with. That is genuinely valuable and it is most of the value.
  • A framework encodes a strategy. Effort in the denominator says "we are capacity-constrained this quarter". Cost of delay in the numerator says "we are revenue-constrained". Both are legitimate. Neither is neutral, and whichever one your team defaults to is a strategy nobody said out loud.
  • A framework does not choose. It cannot, because the weights are yours and the inputs are yours. When a ranking flips under a second lens, the correct reaction is not to refine the inputs until the two agree. It is to say which lens fits this quarter, and why.
The one question that settles it If the spreadsheet had come out the other way, would we have shipped the other thing? If the answer is no, the spreadsheet was decoration, and you can stop paying for the meeting that produces it. If the answer is yes, then the choice of lens deserved more discussion than the weights did.
Self-studyWhy the flip happens so reliably3 min read

It is not a quirk of these three asks. Any small option set will reorder under two framings whenever the options differ on the dimension the framing weights, which is nearly always. Watch what each lens is structurally blind to:

  • Effort-in-the-denominator lenses punish anything expensive, regardless of value. Live notes need a streaming pipeline Cadence does not have, so it loses on cost even when the revenue story is the strongest one in the room.
  • Cost-of-delay lenses punish anything slow to pay back. The 214 support tickets are a real, evidenced, cheap-to-address problem that does not obviously close a deal this quarter, so it drops.
  • Both lenses handle the agent identically. It ranks last under either, because there is no evidence to enter for reach and no engineering estimate for effort. This is the useful part: an unevidenced ask loses under every lens, which is a much better argument to a CEO than your opinion.

The honest summary: frameworks are excellent at killing the unevidenced ask and poor at choosing between two good ones. For the second job you need a stated strategy, which is a leadership artefact and not a spreadsheet output.

Part 2 · covers the inputs frameworks require, and automation bias

Four inputs, and where they came from 10 min live

Every scoring framework needs the same four things: how many people this touches, how sure you are, what it costs to build, and what it costs to wait. Before drafting was free, missing inputs showed up as blanks, and a blank is an honest signal. Now they show up filled in, formatted, and phrased with exactly the same confidence as the one number somebody actually measured. Nothing in the arithmetic distinguishes them, and nothing in the document does either, unless you require it.

Every framework quietly requires these four, and a drafting tool will supply all four What a drafter hands you What it would take to be real Reach a round percentage, with no query behind it a query someone ran, with the date and the filter beside it Confidence 80%, on every row, always, and never with a reason named evidence per claim, or an honestly lower number Effort 2 weeks, with no engineer having seen the ask the engineer who would build it, saying it out loud Cost of delay high / medium / low, which is not a number at all what a month of waiting costs, in churn, deals or ticket load None of the four is a lie. Each is a plausible number with no source, and arithmetic cannot tell the difference. So ask for the source, not the number. One provenance column beside every input: measured, estimated by someone who would know, or invented. Invented inputs are fine to keep. They are not fine to average with measured ones.
🔍 Click to zoom - the same four inputs, arriving with wildly different amounts of truth behind them

The rule is one column, and it costs a PM about four minutes to fill in. Grade every input on this scale, in public, before anything gets averaged.

Provenance gradeWhat it meansWhat it is worth in a rankingCadence example
Measuredsomebody pulled it from a system, and can tell you the date and the filtercan carry a decision on its own; other inputs should be read against itweek-1 transcript revisit rate at 11% of recorded meetings
Estimated by someone who would knowa named person with the relevant scars said it out loud, with a rangegood enough to rank with, as long as the name stays attached to itan engineer's estimate for summaries, given after seeing the ask
Asserted by the requesterthe person who wants it says it is big, and may well be rightworth a follow-up question, never worth a number in a cell"every demo asks for it", offered for live in-meeting notes
Inventeda plausible figure with no source, usually supplied to fill a gapworth nothing arithmetically; keep it as an open question, not a valuethe agent's reach, which nobody has ever tried to size
LiveConfidence is the input that carries all the risk4 min

Of the four, confidence is the one to interrogate, and it is the one nobody interrogates. Two reasons.

  • It multiplies everything else. Reach and impact are guesses about the world; confidence is a guess about your guesses. Set it at 80% across the board and you have quietly declared that the ask with 214 tickets behind it and the ask with nothing behind it are equally believable. That single decision usually matters more than the weights the room spent forty minutes arguing about.
  • It is the cheapest input to fake and the hardest to challenge. Nobody can prove your confidence was wrong until afterwards, and by then the conversation has moved on. So it drifts to whatever number keeps the favoured item at the top, and it does so without anybody being dishonest.

The fix is small and works immediately. Confidence has to be justified in one clause, per row, in words: "80%, because eight interviews and 214 tickets point the same way" versus "80%, because the CEO is sure". Written next to each other, those two do not survive being read aloud, and nobody has to accuse anybody of anything.

The version of this that fails Do not ask for more precision. A request for better numbers gets you the same invented numbers with more decimal places, faster than before, because producing them is now free. Ask for sources. Precision is cheap and provenance is not, so provenance is the thing worth requiring.
Self-studyAutomation bias, in one paragraph you will recognise2 min read

People trust a number more once it has been through a system, and they trust it more again once it has been formatted. This is well documented, it applies to you as much as to your PMs, and drafting tools sit exactly where it does the most damage: they produce output that is well structured, internally consistent, confident in tone, and completely uncoupled from whether anything in it is true.

The practical consequence for a leader is that your instinct is now a bad detector. A ranked table with sourced inputs and a ranked table with invented ones look identical, and the invented one often looks better, because nothing in it is hedged. This is why the provenance column exists as a structural requirement rather than a piece of advice: you cannot rely on noticing, so you make not-noticing impossible.

Part 3 · what you reward, and what you should ask for

Stop rewarding the spreadsheet 6 min live

Here is the uncomfortable part. Prioritisation theatre exists because leaders reward it. A PM who arrives with a scored table and two decimal places looks rigorous, gets approved, and is not asked where the inputs came from. A PM who arrives with three sentences and an honest "we guessed the effort" looks underprepared. As long as that asymmetry holds, your team will keep producing the spreadsheet, and they are responding correctly to what you actually measure.

What to stop rewarding, and what to ask for instead Theatre: a score to two decimals · inputs nobody sourced, averaged into a rank · confidence at 80% on every row, guesses too · effort estimated by a tool, not by a builder · forty minutes spent arguing about weights · the loudest ask re-entered until it wins · nobody can name what the exercise ruled out What to require instead: three lines · the option we chose, and for which segment · the option we rejected, and why it lost · one sentence on what we gave up · plus the date, and whose name is on it Three lines beat a table you cannot audit. a3 covers who signs, and what signing means.
🔍 Click to zoom - the artefact to stop approving, and the one to start asking for
LiveThe three lines, on Cadence3 min

This is what the PM track lands on in b6, written the way you should expect to receive it:

  • Chosen: post-meeting summaries, for team admins on paid workspaces of 5 to 50 seats.
  • Rejected: live in-meeting notes. The research locates the pain after the meeting rather than during it, and live notes need a streaming pipeline that does not exist. The agent is rejected too, for having no evidence at all.
  • Given up: the demo moment sales wanted. Three lost deals were named as evidence and we are choosing not to address them this quarter, knowing sales will raise it again.

Notice what is not there. No score, no weights, no decimal places. Notice also that the third line is the only one that hurts, and it is the line that makes the document a decision rather than a description. If a ranked backlog arrives without an equivalent of that third line, nothing has been decided yet regardless of how much arithmetic is attached.

Self-studyWhen a scoring exercise is worth running anyway2 min read

Three cases where the spreadsheet earns its keep, so this does not become a blanket ban:

  • Twenty-plus options and no shape. Scoring is a good coarse filter for cutting a long list to five. It is a bad instrument for choosing among the final three, which is the exact opposite of how most teams use it.
  • Killing an unevidenced ask. An option with nothing behind it ranks last under every lens, and that is a far better argument in front of a CEO than a personal opinion. The framework's real political value is here.
  • Making a strategy visible. If effort keeps ending up in the denominator, your team believes you are capacity-constrained. Worth knowing, and worth confirming or correcting out loud.

What it is never worth: a 0.3-point margin between two good options on inputs nobody sourced. At that point the score is noise dressed as arithmetic, and someone should just decide and say what they gave up.

Discussion exercise · 12 min · everyone

Score the scores ◆ run this in your next review

Nothing to write and nothing to install. Three prompts, in escalating order of discomfort, each usable in a real product review this week. Run the first one out loud in the room now, on whatever ranked list somebody in this session brought with them.

Somebody produces a real ranked backlog from their own team, from the last month. Not a Cadence example. The exercise only works on a list somebody in the room approved.

Read prompt 1 word for word and grade the top three rows. Time it. Under two minutes means the provenance was already known; over five means it was not.

Count the invented inputs. Not to embarrass anybody: the count is the finding. Most real backlogs come in at half or more, and the number is roughly the same in every company.

Run prompt 2 and see whether the order survives a second lens. If it flips, the room has just watched the framework fail to decide anything.

Close with prompt 3. Whoever brought the list writes the three lines on the spot, out loud. If the third line is hard to say, that is the session working.

LivePrompt 1 · the three-word grading4 min
Say this, word for word "Take the top three items on this list. For each input in the score - reach, confidence, effort, cost of delay - give me one of three words: measured, estimated, or invented. You do not need to defend the numbers, and nothing bad happens to an invented one. I just want the twelve words."

What a good answer sounds like: a mix, delivered without defensiveness. "Reach is measured, I pulled it Tuesday. Effort is estimated, Priya gave me three weeks after seeing the ticket sample. Cost of delay is invented, I have no idea what a month of waiting costs us." A strong PM will volunteer the invented ones first, because they already knew, and will often tell you what it would take to upgrade one.

The failure mode to listen for: everything comes back "measured". That is not a rigorous team, it is a team that has not looked. Second failure: the PM starts defending the ranking instead of grading the inputs, which means the question landed as a challenge to their judgement rather than a question about sources, so say the last sentence of the prompt again. Third: confidence is 80% on every row and nobody can say why, which is the single most common finding in this exercise.

LivePrompt 2 · the second lens4 min
Say this, word for word "Which lens is this, and why does that lens fit this quarter? Then rank the same three under one that weights the opposite thing, and tell me whether the order changes."

What a good answer sounds like: the PM names the framing and connects it to a real constraint. "Effort is in the denominator because we lost two engineers and everything expensive is off the table until March." Then, when the order flips, they say which lens they trust and why, rather than trying to reconcile the two. Best possible answer: "it flips, and I think the second lens is the honest one, which means I have been recommending the wrong thing."

The failure mode to listen for: "we always use RICE." That is a strategy nobody chose, inherited from a blog post, and it is quietly deciding your roadmap. Also listen for an attempt to tune the inputs until both lenses agree. That is not analysis, it is fitting the numbers to a conclusion, and it usually happens in good faith, which is what makes it hard to spot.

Self-studyPrompt 3 · the three lines, from memory4 min
Say this, word for word "Forget the spreadsheet for a second. Three lines: what we are doing and for whom, what we are not doing and why it lost, and one sentence on what we gave up. If the third line is blank, we have not decided yet."

What a good answer sounds like: specific on all three, and slightly painful on the third. "We are doing summaries for paid team admins. We are not doing live notes, because the pain is after the meeting and we have no streaming pipeline. We are giving up the demo moment sales asked for, and three lost deals stay unaddressed this quarter." The third line names a person who will be unhappy. That is what makes it a real trade-off rather than a summary.

The failure mode to listen for: the third line comes back as a benefit in disguise. "We gave up scope creep" and "we gave up complexity" are not trade-offs, they are compliments. If nothing was surrendered and nobody is worse off, no choice was made: the options were not really in competition, or the answer was obvious and did not need a meeting. Both are worth knowing.

Real world

The tell is always confidence, and it is always 80%. Run prompt 1 on any real scored backlog and you will find the same artefact: five rows, five different reach numbers, five different effort estimates, and confidence sitting at 80% on every single line including the one with no evidence behind it at all. Nobody chose that. It happens because confidence is the only input that feels like an opinion rather than a fact, so it gets set once, early, at a number that sounds appropriately humble, and then never revisited. The consequence is that the input carrying the most risk in the whole calculation is the one input that was never actually filled in, and the ranking silently treats a well-evidenced ask and a CEO's hunch as equally believable. You can find this in ten seconds by scanning one column, which makes it the best value question a leader can ask in a prioritisation review.

Take this to your next review

The questions to ask your PMs ◐ 7 questions

Each of these is short, none of them reads as an attack, and every one of them is answerable in under a minute by a PM who did the work. That last property is what makes them useful: the cost of asking is trivial, and the information in a hesitation is high.

Homework

Try it yourself - this week ◐ 30-40 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

Prioritisation frameworks (RICE, WSJF, cost of delay) as lenses rather than answersPart 1 · the same three asks, ranked twice, two orders
The inputs every framework quietly requires, and the provenance gradesPart 2 · the four inputs and the one column that fixes them
AI-assisted knowledge work: automation bias and confident unsourced numbersPart 2 self-study · why your instinct is now a bad detector
Spec practice: the explicit trade-off as the line that makes a document a decisionPart 3 · taught in full in a3, and scored in the PM track's b5
Metric trees and driver decomposition behind a reach or impact numberOut of scope by design - learn-metric-decomposition
Estimation technique, sequencing and capacity planningOut of scope by design - learn-ai-project-management
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · The same three options rank in two different orders under two different frameworks. What should you conclude?

Refining inputs until two lenses agree is fitting numbers to a conclusion. The productive move is to say which lens fits this quarter and why, because that framing choice is the actual decision and it is made by a person before any arithmetic runs.

2 · A PM brings a scored backlog with confidence at 80% on all five rows, including one ask with no evidence behind it. What is your best first move?

More precision gets you the same invented numbers with more decimal places, because producing them is free. Provenance is the expensive thing, so provenance is what to require. A uniform confidence column means the input carrying all the risk was never actually filled in.

3 · Your team's ranking puts option A ahead of option B by 0.3 points, on inputs that are mostly estimates. What do you require before it becomes a plan?

A 0.3-point margin on unsourced inputs is noise wearing arithmetic. Three lines can be audited a year later, name who is worse off, and force somebody to have actually chosen. The spreadsheet is optional; those lines are not.

Leader session 4 cheat sheet · pin this

Frameworks are lensesSame three options, two framings, two orders. The framework made a decision legible; it did not make one.
Choosing the lens is the decisionEffort in the denominator says capacity-constrained. Cost of delay in the numerator says revenue-constrained. Neither is neutral.
Four inputs, alwaysReach, confidence, effort, cost of delay. A drafting tool supplies all four in the same confident register.
The provenance columnMeasured · estimated by someone who would know · asserted by the requester · invented. Four minutes to fill in.
Confidence carries the riskIt multiplies everything and gets faked most. 80% on every row means it was never filled in.
Ask for sources, not precisionPrecision is free now. Provenance is not, so provenance is the thing worth requiring.
The three linesChosen and for whom · rejected and why · what we gave up. The third one is the only one that hurts.
Where this goes nextThe PM track's b6 makes this exact call on Cadence: summaries, at the cost of the demo moment sales wanted. Next: a5, metrics.