Why this session exists
You have sat in the meeting. A PM presents a scored backlog, the scores run to two decimal places, the top item wins by a fraction, and everybody nods because the process was followed. Then somebody asks where the reach number came from and the room goes quiet. That silence is the subject of this session, and it got louder the moment drafting became free: a tool will now fill in reach, confidence, effort and cost of delay for you, instantly, in the same confident register as the numbers that came out of your warehouse.
Session a1 set the production and commitment split, a2 set the evidence standard, and a3 covered what a spec has to contain before anyone signs it. This session is the ranking meeting: how to read a score without being captured by it, what to demand about where each input came from, and what to accept as the record of the decision. It ends where every session in this track ends, with the questions to ask your PMs.
The same options, ranked twice 11 min live
Cadence has three asks on the table: live in-meeting notes, which sales wants because every demo asks for it and three deals were lost; post-meeting summaries, which support wants because 214 tickets in three months are tagged "cannot find what was decided"; and an agent, which the CEO wants and for which no evidence has been offered. Below, one PM ranks those three twice in an afternoon. Both rankings are competent. Both are arithmetically correct. They disagree about what to build.
LiveWhat a framework is actually for4 min▶
None of this makes scoring useless. It makes scoring a different thing than most teams present it as, and the difference matters to you specifically, because you are the one being asked to approve the output.
- A framework makes a decision legible. It forces every option into the same units, so the comparison can be argued with. That is genuinely valuable and it is most of the value.
- A framework encodes a strategy. Effort in the denominator says "we are capacity-constrained this quarter". Cost of delay in the numerator says "we are revenue-constrained". Both are legitimate. Neither is neutral, and whichever one your team defaults to is a strategy nobody said out loud.
- A framework does not choose. It cannot, because the weights are yours and the inputs are yours. When a ranking flips under a second lens, the correct reaction is not to refine the inputs until the two agree. It is to say which lens fits this quarter, and why.
Self-studyWhy the flip happens so reliably3 min read▶
It is not a quirk of these three asks. Any small option set will reorder under two framings whenever the options differ on the dimension the framing weights, which is nearly always. Watch what each lens is structurally blind to:
- Effort-in-the-denominator lenses punish anything expensive, regardless of value. Live notes need a streaming pipeline Cadence does not have, so it loses on cost even when the revenue story is the strongest one in the room.
- Cost-of-delay lenses punish anything slow to pay back. The 214 support tickets are a real, evidenced, cheap-to-address problem that does not obviously close a deal this quarter, so it drops.
- Both lenses handle the agent identically. It ranks last under either, because there is no evidence to enter for reach and no engineering estimate for effort. This is the useful part: an unevidenced ask loses under every lens, which is a much better argument to a CEO than your opinion.
The honest summary: frameworks are excellent at killing the unevidenced ask and poor at choosing between two good ones. For the second job you need a stated strategy, which is a leadership artefact and not a spreadsheet output.
Four inputs, and where they came from 10 min live
Every scoring framework needs the same four things: how many people this touches, how sure you are, what it costs to build, and what it costs to wait. Before drafting was free, missing inputs showed up as blanks, and a blank is an honest signal. Now they show up filled in, formatted, and phrased with exactly the same confidence as the one number somebody actually measured. Nothing in the arithmetic distinguishes them, and nothing in the document does either, unless you require it.
The rule is one column, and it costs a PM about four minutes to fill in. Grade every input on this scale, in public, before anything gets averaged.
| Provenance grade | What it means | What it is worth in a ranking | Cadence example |
|---|---|---|---|
| Measured | somebody pulled it from a system, and can tell you the date and the filter | can carry a decision on its own; other inputs should be read against it | week-1 transcript revisit rate at 11% of recorded meetings |
| Estimated by someone who would know | a named person with the relevant scars said it out loud, with a range | good enough to rank with, as long as the name stays attached to it | an engineer's estimate for summaries, given after seeing the ask |
| Asserted by the requester | the person who wants it says it is big, and may well be right | worth a follow-up question, never worth a number in a cell | "every demo asks for it", offered for live in-meeting notes |
| Invented | a plausible figure with no source, usually supplied to fill a gap | worth nothing arithmetically; keep it as an open question, not a value | the agent's reach, which nobody has ever tried to size |
LiveConfidence is the input that carries all the risk4 min▶
Of the four, confidence is the one to interrogate, and it is the one nobody interrogates. Two reasons.
- It multiplies everything else. Reach and impact are guesses about the world; confidence is a guess about your guesses. Set it at 80% across the board and you have quietly declared that the ask with 214 tickets behind it and the ask with nothing behind it are equally believable. That single decision usually matters more than the weights the room spent forty minutes arguing about.
- It is the cheapest input to fake and the hardest to challenge. Nobody can prove your confidence was wrong until afterwards, and by then the conversation has moved on. So it drifts to whatever number keeps the favoured item at the top, and it does so without anybody being dishonest.
The fix is small and works immediately. Confidence has to be justified in one clause, per row, in words: "80%, because eight interviews and 214 tickets point the same way" versus "80%, because the CEO is sure". Written next to each other, those two do not survive being read aloud, and nobody has to accuse anybody of anything.
Self-studyAutomation bias, in one paragraph you will recognise2 min read▶
People trust a number more once it has been through a system, and they trust it more again once it has been formatted. This is well documented, it applies to you as much as to your PMs, and drafting tools sit exactly where it does the most damage: they produce output that is well structured, internally consistent, confident in tone, and completely uncoupled from whether anything in it is true.
The practical consequence for a leader is that your instinct is now a bad detector. A ranked table with sourced inputs and a ranked table with invented ones look identical, and the invented one often looks better, because nothing in it is hedged. This is why the provenance column exists as a structural requirement rather than a piece of advice: you cannot rely on noticing, so you make not-noticing impossible.
Stop rewarding the spreadsheet 6 min live
Here is the uncomfortable part. Prioritisation theatre exists because leaders reward it. A PM who arrives with a scored table and two decimal places looks rigorous, gets approved, and is not asked where the inputs came from. A PM who arrives with three sentences and an honest "we guessed the effort" looks underprepared. As long as that asymmetry holds, your team will keep producing the spreadsheet, and they are responding correctly to what you actually measure.
LiveThe three lines, on Cadence3 min▶
This is what the PM track lands on in b6, written the way you should expect to receive it:
- Chosen: post-meeting summaries, for team admins on paid workspaces of 5 to 50 seats.
- Rejected: live in-meeting notes. The research locates the pain after the meeting rather than during it, and live notes need a streaming pipeline that does not exist. The agent is rejected too, for having no evidence at all.
- Given up: the demo moment sales wanted. Three lost deals were named as evidence and we are choosing not to address them this quarter, knowing sales will raise it again.
Notice what is not there. No score, no weights, no decimal places. Notice also that the third line is the only one that hurts, and it is the line that makes the document a decision rather than a description. If a ranked backlog arrives without an equivalent of that third line, nothing has been decided yet regardless of how much arithmetic is attached.
Self-studyWhen a scoring exercise is worth running anyway2 min read▶
Three cases where the spreadsheet earns its keep, so this does not become a blanket ban:
- Twenty-plus options and no shape. Scoring is a good coarse filter for cutting a long list to five. It is a bad instrument for choosing among the final three, which is the exact opposite of how most teams use it.
- Killing an unevidenced ask. An option with nothing behind it ranks last under every lens, and that is a far better argument in front of a CEO than a personal opinion. The framework's real political value is here.
- Making a strategy visible. If effort keeps ending up in the denominator, your team believes you are capacity-constrained. Worth knowing, and worth confirming or correcting out loud.
What it is never worth: a 0.3-point margin between two good options on inputs nobody sourced. At that point the score is noise dressed as arithmetic, and someone should just decide and say what they gave up.
Score the scores ◆ run this in your next review
Nothing to write and nothing to install. Three prompts, in escalating order of discomfort, each usable in a real product review this week. Run the first one out loud in the room now, on whatever ranked list somebody in this session brought with them.
Somebody produces a real ranked backlog from their own team, from the last month. Not a Cadence example. The exercise only works on a list somebody in the room approved.
Read prompt 1 word for word and grade the top three rows. Time it. Under two minutes means the provenance was already known; over five means it was not.
Count the invented inputs. Not to embarrass anybody: the count is the finding. Most real backlogs come in at half or more, and the number is roughly the same in every company.
Run prompt 2 and see whether the order survives a second lens. If it flips, the room has just watched the framework fail to decide anything.
Close with prompt 3. Whoever brought the list writes the three lines on the spot, out loud. If the third line is hard to say, that is the session working.
LivePrompt 1 · the three-word grading4 min▶
What a good answer sounds like: a mix, delivered without defensiveness. "Reach is measured, I pulled it Tuesday. Effort is estimated, Priya gave me three weeks after seeing the ticket sample. Cost of delay is invented, I have no idea what a month of waiting costs us." A strong PM will volunteer the invented ones first, because they already knew, and will often tell you what it would take to upgrade one.
The failure mode to listen for: everything comes back "measured". That is not a rigorous team, it is a team that has not looked. Second failure: the PM starts defending the ranking instead of grading the inputs, which means the question landed as a challenge to their judgement rather than a question about sources, so say the last sentence of the prompt again. Third: confidence is 80% on every row and nobody can say why, which is the single most common finding in this exercise.
LivePrompt 2 · the second lens4 min▶
What a good answer sounds like: the PM names the framing and connects it to a real constraint. "Effort is in the denominator because we lost two engineers and everything expensive is off the table until March." Then, when the order flips, they say which lens they trust and why, rather than trying to reconcile the two. Best possible answer: "it flips, and I think the second lens is the honest one, which means I have been recommending the wrong thing."
The failure mode to listen for: "we always use RICE." That is a strategy nobody chose, inherited from a blog post, and it is quietly deciding your roadmap. Also listen for an attempt to tune the inputs until both lenses agree. That is not analysis, it is fitting the numbers to a conclusion, and it usually happens in good faith, which is what makes it hard to spot.
Self-studyPrompt 3 · the three lines, from memory4 min▶
What a good answer sounds like: specific on all three, and slightly painful on the third. "We are doing summaries for paid team admins. We are not doing live notes, because the pain is after the meeting and we have no streaming pipeline. We are giving up the demo moment sales asked for, and three lost deals stay unaddressed this quarter." The third line names a person who will be unhappy. That is what makes it a real trade-off rather than a summary.
The failure mode to listen for: the third line comes back as a benefit in disguise. "We gave up scope creep" and "we gave up complexity" are not trade-offs, they are compliments. If nothing was surrendered and nobody is worse off, no choice was made: the options were not really in competition, or the answer was obvious and did not need a meeting. Both are worth knowing.
The tell is always confidence, and it is always 80%. Run prompt 1 on any real scored backlog and you will find the same artefact: five rows, five different reach numbers, five different effort estimates, and confidence sitting at 80% on every single line including the one with no evidence behind it at all. Nobody chose that. It happens because confidence is the only input that feels like an opinion rather than a fact, so it gets set once, early, at a number that sounds appropriately humble, and then never revisited. The consequence is that the input carrying the most risk in the whole calculation is the one input that was never actually filled in, and the ranking silently treats a well-evidenced ask and a CEO's hunch as equally believable. You can find this in ten seconds by scanning one column, which makes it the best value question a leader can ask in a prioritisation review.
The questions to ask your PMs ◐ 7 questions
Each of these is short, none of them reads as an attack, and every one of them is answerable in under a minute by a PM who did the work. That last property is what makes them useful: the cost of asking is trivial, and the information in a hesitation is high.
- "Which of these inputs is measured, which is estimated, and which is invented?" Grades the score without disputing it. The answer tells you exactly how much weight the ranking can carry, and it is the only question on this list you should ask every single time.
- "Why 80% and not 40%?" Confidence multiplies everything and is the input most often set by habit. You will get either a sentence of evidence or a pause, and both are useful.
- "Who gave you the effort number, and had they seen the ask?" An estimate with no builder behind it is a wish. A name makes it checkable, and the person named usually revises it downward or upward within a day.
- "What would have to be true for the second-ranked option to win?" Tests whether the gap is real. If a small change to one unsourced input reorders the list, the margin is noise and the decision needs an argument rather than arithmetic.
- "Which lens is this, and why that lens this quarter?" Makes the framing visible as a choice, which is where the deciding actually happened. Teams that cannot answer are running a strategy they inherited rather than chose.
- "What did we give up by choosing this?" The single most useful question in a prioritisation review, and the same test a3 applied to specs. A blank answer means the options were never really in competition.
- "If the numbers had come out the other way, would we ship the other thing?" Ask this once a quarter. A no tells you the exercise is ceremony, and you can spend the meeting on the argument instead.
Try it yourself - this week ◐ 30-40 min total
- Take the last ranked backlog you approved and add the provenance column yourself, from memory. Count the invented rows before you ask anyone else to. If you cannot grade your own approved decision, that is the finding.
- Ask one PM the three-word question in your next review, and note how long the answer takes. Time is the measure here, not the answer.
- Choose the lens your team will default to this quarter, say why in one sentence, and tell them it was your choice rather than a property of the universe.
- Find one decision from last quarter with no rejected option written down anywhere, and write it now from memory, while you still have it. Next quarter you will not.
- Cancel one weight-tuning meeting and replace it with the three lines. Watch whether decision quality drops. It will not, and the hour is worth having back.
Sources covered
Full source map in materials/official-course-map.md. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · The same three options rank in two different orders under two different frameworks. What should you conclude?
Refining inputs until two lenses agree is fitting numbers to a conclusion. The productive move is to say which lens fits this quarter and why, because that framing choice is the actual decision and it is made by a person before any arithmetic runs.
2 · A PM brings a scored backlog with confidence at 80% on all five rows, including one ask with no evidence behind it. What is your best first move?
More precision gets you the same invented numbers with more decimal places, because producing them is free. Provenance is the expensive thing, so provenance is what to require. A uniform confidence column means the input carrying all the risk was never actually filled in.
3 · Your team's ranking puts option A ahead of option B by 0.3 points, on inputs that are mostly estimates. What do you require before it becomes a plan?
A 0.3-point margin on unsourced inputs is noise wearing arithmetic. Three lines can be audited a year later, name who is worse off, and force somebody to have actually chosen. The spreadsheet is optional; those lines are not.