Frameworks do not decide anything
By now the three asks are comparable. b2 gave each one an evidenced problem, b3 verified the research behind them, b4 named the job and wrote down the branches we cut, and b5 turned the leading candidate into something signable. Everything since b1 has been about making the asks sit in the same units. This session spends that work: two scoring lenses, one decision, and the sentence that says what the decision cost.
The reason prioritisation goes wrong is not that teams pick the wrong framework. It is that a framework returns a number, a number looks like an answer, and nobody asks where the inputs came from. AI makes this worse in a very specific way: ask it to score a backlog and it will fill in reach, impact, confidence and effort for all three items, in seconds, with two decimal places and no basis whatsoever. The output is arithmetic on guesses, and arithmetic is extremely convincing.
A framework is a lens, not an answer 7 min live
Two lenses, same three asks, same evidence, no cheating: RICE, which divides expected value by effort, and WSJF, which divides cost of delay by how long the job takes. Both are respectable. Both are used by serious teams. And they disagree, which is the useful part, because the disagreement tells you exactly what each lens is willing to ignore.
LiveWhat each lens is willing to ignore4 min▶
Read the two formulas as statements of belief, because that is what they are:
- RICE believes cheap reach wins. Effort is a divisor, so a 3-week job with broad reach beats an 11-week job almost regardless of how valuable the 11-week job is. It also has an explicit confidence term, which means it is the only one of the two that can express "we are not sure". That single term is why the agent finishes 31x behind rather than second.
- WSJF believes urgency wins. Cost of delay rewards anything framed as time-critical or strategically enabling. It has no confidence input at all. So an ask with zero evidence and a strong narrative scores exactly as well as an ask with 214 tickets behind it, provided somebody says the words "market window".
- Which is why lens choice is a decision, not a preliminary. Whoever picks the framework has already narrowed the answer. If someone arrives at a review with a ranking, the first question is not "what did it score" but "which lens, and who chose it".
Self-studyWhy the score is not the artifact3 min read▶
Three things a score does well, and one it cannot do at all.
- It forces the inputs into the open. You cannot score an ask without stating how many people it reaches and how long it takes, and half the value of the exercise is watching somebody discover they cannot state either.
- It makes disagreement specific. "I think this is more important" becomes "I think your confidence of 0.5 should be 0.8, and here is why", which is a conversation that can actually finish.
- It kills the bottom of the list cheaply. Items that lose under every lens are genuinely dead, and that is worth knowing early.
- It cannot choose. A number ranks; a person commits. The gap between the two is the entire second half of this session, and no amount of scoring closes it.
The leader-track version of this is a4, prioritisation without theatre, which is about spotting a ranking built on guesses in somebody else's deck rather than building one yourself.
Every score needs a provenance column 6 min live
Both formulas above quietly demand four things: reach, confidence, effort, and some notion of what waiting costs. None of the four arrives with the ask. All four have to be produced by somebody, and if you do not produce them, a drafting tool will, in a tone of complete assurance. The fix is small and it changes everything: one extra column that says whether each number was measured, estimated by someone who would know, or invented.
| Input, as used on this page | Value | Class | What would move it up a class |
|---|---|---|---|
| Reach · summaries | 100 of 100 recorded meetings | Measured | nothing needed: every recorded meeting on a paid workspace gets one, and analytics counts recorded meetings already |
| Reach · live notes | 68 of 100 meetings | Measured, self-reported | the 68% comes from a survey answer, not observed behaviour; instrument whether notes are taken elsewhere during a call |
| Reach · the agent | 30 of 100 meetings | Invented | one interview in which an admin says what an agent would do for them, and how often |
| Impact · summaries | 2 of 3 | Estimated by someone who would know | the support lead scored it against 214 tickets and 6 of 8 interviews; a ticket-to-account rollup would tighten it |
| Impact · live notes | 3 of 3 | Invented | pull the loss reasons on the 3 lost deals and check whether live notes was actually named; until then this is an assertion by an interested party |
| Confidence · summaries | 0.8 | Estimated by someone who would know | two independent signals agreeing (tickets and interviews) is what buys the 0.8; a third would not add much |
| Confidence · live notes | 0.5 | Estimated by someone who would know | one unverified signal, so half. Verify the deals and this moves in either direction, which is the point |
| Confidence · the agent | 0.25 | Estimated against nothing | the number a drafting tool writes here is 0.8, and the number the room says out loud is 0.8, and neither has anything under it |
| Effort · all three | 3 / 11 / 13 weeks | Estimated by someone who would know | all three are engineering estimates and labelled as estimates on this page; the 11 includes a streaming pipeline that does not exist |
| Run cost · summaries | 0.02 per meeting | Measured | nothing needed, and worth noticing that it is the only cost number on the board with a real basis |
LiveConfidence is where the risk lives, and where the lying happens4 min▶
Of the four inputs, confidence is the strange one. Reach and effort are wrong by a factor of two at worst. Confidence is a multiplier on your entire belief that the thing works at all, and unlike the others it has no natural units, which makes it the easiest number in the whole exercise to move without anyone noticing.
- It absorbs every unresolved question. Did the research cover the right segment? Was the sample friendly? Is the mechanism plausible? All of it lands in one decimal, and then the decimal gets multiplied by a big reach number and disappears.
- It moves to fit the answer people want. Nobody writes 0.25 next to the CEO's request. Watch what happens to a confidence value between the draft and the review, and who was in the room when it changed.
- A drafting tool always writes it high. Asked to score a backlog it produces confidence values in the 0.7 to 0.9 band across the board, because high confidence is what confident prose sounds like. It has not seen your research. It has not seen anything.
- So state the basis, not the number. "0.8, two independent signals" and "0.8, because it feels likely" are the same cell and completely different claims. The provenance column is what makes them look different on the page.
LiveWhat to hand a drafting tool, and what to never let it touch3 min▶
There is real work for AI in this session, and it is all on one side of the line.
- Genuinely useful: converting your inputs into both scorings without arithmetic slips, generating the sensitivity table (which single input, changed by how much, flips the ranking), writing the strongest possible case for the option you are about to cut, and drafting the note to sales once you have decided what it says.
- Actively harmful: supplying any of the four inputs. Ask for a scored backlog and you get one, complete, plausible and untraceable, and the provenance column comes back full of numbers that belong in the invented row without saying so.
- The tell: if every confidence value in a scored backlog sits between 0.7 and 0.9, nobody scored anything. Real confidence values are lumpy, because real evidence is lumpy.
Then somebody has to choose 6 min live
The two lenses agree on the winner and disagree on everything else, so the honest reading is: summaries survived both framings, and the ordering below it is noise. That is as far as arithmetic goes. What follows is a choice, made by a person, on evidence, with a cost attached. Here it is, made on the page rather than left as an exercise.
LiveWhy summaries, in the four sentences that actually did the work4 min▶
Neither score is the reason. These are:
- The evidence locates the pain after the meeting. 214 tickets tagged "cannot find what was decided" and 6 of 8 interviewed admins who never reopen a transcript describe the same moment, and it is not the moment live notes serves. This is the argument. Everything else is supporting detail.
- The two strongest signals are independent. Support tags and research interviews are different instruments pointed at different people, and they agree. That is what a confidence of 0.8 is actually made of.
- It is the only option we can be wrong about cheaply. 3 weeks and 0.02 a meeting means being wrong costs a month. 11 weeks plus a pipeline that does not exist means being wrong costs the quarter and leaves infrastructure nobody asked for.
- And it makes the live-notes question answerable sooner. Shipping summaries first buys eight weeks in which sales can pull the loss reasons on those three deals. If live notes is genuinely deal-blocking, that argument comes back stronger next cycle with a number in it.
Self-studySaying no to the CEO's agent without saying no to the CEO3 min read▶
The agent is the hardest of the three to cut, and not because it scores badly. It scores badly under both lenses. It is hard to cut because of who asked, which is a political problem wearing a prioritisation costume.
What does not work: presenting the score. A number that says the CEO's idea ranks third reads as a rebuttal, and the reply is that the framework is wrong, which is a debate you will lose because the framework is partly wrong, as this page just demonstrated.
What does work: convert the no into a missing input. "There is no version of this I can size, because nobody has told me which job it does for which admin. Give me one customer conversation where an agent is the obvious answer, and I will bring it back scored." That is not a refusal, it is a request, and the request is genuinely the blocker. If the conversation never happens, the ask dies of natural causes and nobody had to fight about it. If it does happen, you were wrong and now you have the input you were missing.
The stakeholder-facing version of this conversation, including the three questions the CEO will ask next, is b9.
The scored board and the decision record ◆ real material to judge
Both scorings on one board, with the provenance of the weakest input in the last column, because that column is the one that decides how much of the row you are allowed to believe. Read the rank-move column before the score columns.
Write the inputs down before you score anything. Reach with its denominator, impact on a stated scale, confidence as a decimal, effort in weeks from the people who would do the work. If a cell is empty, leave it empty and see who volunteers a number.
Label every cell: measured, estimated by someone who would know, or invented. Do this before the arithmetic, not after, or the labels will drift towards whatever justifies the ranking you already have.
Score under two lenses. Then compare the rankings, and more importantly compare the margins. A winner that holds under both lenses but goes from 31x clear to 1.4x clear is telling you the gap is a modelling artifact.
Find the single input carrying the result. Change one cell at a time and watch the order. Here it is effort: move the summaries estimate from 3 weeks to 9 and the RICE score falls from 53.3 to 17.8, still first. Move it to 30 and it loses. So the decision survives a 3x estimation error and not a 10x one, which is worth knowing.
Then stop scoring and write the record. The decision, the reason, what it cost, what you will say to the person who lost, the cut branches, and the condition that re-opens each one. This artifact outlives the spreadsheet by about a year.
| Ask | RICE inputs R · I · C ÷ E | RICE | WSJF inputs BV + TC + OE ÷ weeks | WSJF | Rank move | Weakest input, and where that number came from |
|---|---|---|---|---|---|---|
| Post-meeting summaries | 100 · 2 · 0.8 ÷ 3 | 53.3 | 13 + 3 + 3 ÷ 3 | 6.3 | 1 → 1 | Effort, 3 weeks: estimated by the tech lead who would build it. The 0.02 per meeting run cost is measured, which makes this the best-sourced row on the board |
| Live in-meeting notes | 68 · 3 · 0.5 ÷ 11 | 9.3 | 20 + 20 + 8 ÷ 11 | 4.4 | 2 → 3 | Impact, scored a generous 3 of 3: invented. It rests entirely on 3 lost deals that nobody has checked against the loss reasons, asserted by the team that wants the feature |
| The agent | 30 · 3 · 0.25 ÷ 13 | 1.7 | 20 + 20 + 20 ÷ 13 | 4.6 | 3 → 2 | Reach, 30 of 100: invented, and so are impact and all three cost-of-delay components. Five invented cells producing a second-place finish under lens B |
Note what the WSJF column did to the bottom of the board. The ask with the most invented inputs finished second, ahead of the ask with one unverified signal, because the lens has no confidence term and rewards asserted urgency. Nothing dishonest happened. Somebody just chose a framework.
DECISION RECORD · Cadence · next quarter · owner: PM · reviewed with: eng lead, support lead, sales lead
WE ARE BUILDING Post-meeting summaries, for team admins on paid workspaces of
5 to 50 seats. Signable spec: see b5, scored 100 of 100.
BECAUSE The evidence locates the pain AFTER the meeting rather than during
it. 214 support tickets in 3 months tagged "cannot find what was
decided", the second most common tag; and 6 of 8 interviewed admins
never reopen a transcript at all. Two independent instruments,
same finding.
AT THE COST OF Live in-meeting notes. This ships without the demo moment sales
asked for, and we are accepting that for at least one quarter.
WHAT I WILL SAY "Your three lost deals are the strongest thing anybody brought me,
TO SALES and I could not check a single one of them. So here is the deal:
pull the loss reasons on those three, and flag any deal this quarter
where live notes gets named, and I will re-score it next cycle with
your number instead of my guess. Meanwhile the thing 214 tickets are
about takes 3 weeks rather than 11, which means your argument gets
its hearing sooner this way than if I started live notes today."
CUT, AND WHY Live in-meeting notes · impact rests on 3 unverified deals, and it
needs a streaming pipeline that does not exist. 11 weeks.
The agent · no evidence was offered, and no job has been named.
13 weeks at the low end.
RE-OPEN IF The loss-reason pull names live notes in 3 or more deals; or an admin
conversation names a specific job an agent does; or the summaries
metric set in b7 does not move.
WHAT WOULD HAVE A verified deal-loss number for live notes. Reach and impact for the
CHANGED THIS agent that are not invented. A summaries effort estimate above
30 weeks.
Lens shopping is the most common form of prioritisation theatre, and it never looks like cheating. A team runs RICE, does not like the answer, and someone points out that RICE undervalues strategic bets. Which is true. So they run WSJF, and the strategic bet comes second instead of last, and the deck says "scored under two frameworks" as though that were rigour rather than the opposite. The countermeasure is not banning frameworks. It is fixing the lens and the provenance labels before anybody sees a ranking, and writing down which lens you chose and why in the same document as the score. A ranking is only evidence if the method was picked before the answer was known.
Your turn: three passes over the board ◆ 8 min
No tool needed. A pen and the table above.
LiveQ1 · Re-score the agent with honest confidence3 min▶
The agent's WSJF row is 20 + 20 + 20 ÷ 13 = 4.6, second place. Every one of those three 20s is invented, and WSJF has no confidence term to catch that. So add one by hand: multiply the cost of delay by the confidence you would actually defend, which for an ask with no evidence at all is 0.25.
60 x 0.25 = 15, over 13 weeks, gives 1.2. Now it is last under both lenses, by a wide margin, and it got there without anyone arguing about whether agents matter. That is the move: when a framework has no slot for how much you believe the inputs, add the slot rather than accepting the ranking. Then say so in the document, because you have just modified the method and that is a thing reviewers are entitled to know.
LiveQ2 · Find the input doing all the work3 min▶
Take the summaries row and change one input at a time until the ranking breaks. Then do the same for live notes.
You will find two things. First, the summaries decision is robust to reach and impact being wrong by a lot, and fragile to the effort estimate being wrong by a factor of ten, which is unlikely for a 3-week job and very likely for an 11-week one. Second, live notes flips ahead of summaries under RICE if its confidence goes from 0.5 to 0.9 and its effort comes in at 4 weeks rather than 11. So the whole live-notes case reduces to two claims: the deals are real, and the pipeline is easier than we think. Both are checkable, neither has been checked, and now they are written down as the two things that would change the decision.
This is the sensitivity pass, and it is the one part of the session worth handing to a tool: give it your inputs and ask which single cell, changed by how much, flips the order. It is arithmetic on numbers you supplied, which is exactly the safe half of the line.
Self-studyQ3 · Audit a ranking you did not build4 min read▶
Find the most recent prioritised backlog in your company that you did not personally score. Add one column: measured, estimated by someone who would know, or invented. Do not change any numbers and do not tell anyone yet.
Two patterns show up almost every time. The confidence column is uniformly high, which means nobody scored confidence, they scored enthusiasm. And the items at the top with the smallest effort estimates were estimated by the person who wanted them built, not by the person who would build them. Neither finding requires a confrontation. Both can be raised as one question in the next planning meeting: "can we mark which of these numbers we measured?" It reads as tidiness, it is answerable, and the answer reorders the list.
Try it yourself - this week ◐ 25-35 min total
- Score your carried backlog item under two lenses, not one. Use RICE and either WSJF or a plain cost-of-delay reading, and write both rankings down even when they agree.
- Add the provenance column and fill it in honestly: measured, estimated by someone who would know, or invented. Count the invented cells. That count is the finding, not the score.
- Run the sensitivity pass. Name the single input that, if wrong, changes your answer, and write one line on how you would check it and what it would cost to check.
- Write the trade-off sentence: what you are building, what you are therefore not building, who loses, and the words you will say to them. If you cannot write the last part, the decision is not made.
- Write the cut record: each rejected option, why, and the condition that re-opens it. Put it where the argument will happen again, which is usually the top of the planning doc rather than a folder nobody opens.
Sources covered
Full source map in materials/official-course-map.md. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · The agent finishes third under RICE (1.7) and second under WSJF (4.6). What changed?
RICE multiplies by confidence, which is the only place either formula can record "we have nothing behind this". Remove that term and five invented cells produce a second-place finish. Neither framework is wrong; they disagree about what is allowed to be ignored, and that disagreement is the finding.
2 · You need a reach number for an ask and nobody has one. What do you do?
The problem is never that a number is missing, it is that a missing number gets replaced silently and then survives into a plan. Labelling it invented keeps the row on the board, keeps the gap visible, and turns the next step into a small research task rather than an argument about priorities.
3 · Summaries win under both lenses. Has the framework made the decision?
A number ranks; a person commits. The test for whether you decided rather than ranked is whether you can name what you gave up, who loses by it, and what you will say to them. Nothing in either formula produces any of those three.