The closing session, and the one with money in it
a1 drew the line between production and commitment. a2 set the evidence standard, a3 named what a spec must contain before it is signable, a4 showed a ranking collapsing under a second lens, and a5 asked whether your organisation can hear a no. Every one of those was a judgement you make in a room you are already sitting in. This one is different: it asks what you buy, in what order, and what you take off your PMs' performance reviews to make any of it stick.
The reason the order matters more than the total is simple arithmetic about bottlenecks. Documents used to be expensive and judgement was the free part. That has inverted. Your organisation can now produce more specs, more research summaries, more competitive analyses and more roadmap variants than it has ever been able to read properly, and the constraint has moved to the only place it could move: the people who decide whether any of it is true. Fund in the wrong order and you will spend real money making the bottleneck worse.
Fund in this order, and the order is the argument 8 min live
Four rungs. Most organisations fund them in exactly the reverse sequence, because tooling is the only one that arrives with an invoice, a vendor, and a slide. The other three are people, habits and unglamorous analyst afternoons, which means nobody proposes them and nothing bad happens for a quarter. Then the readouts stop being useful and nobody can say precisely when that started.
| Order | Fund this | What the money actually buys | Why it sits at this rung and not lower |
|---|---|---|---|
| 1 | Review capacity | Senior time to read properly, a named reviewer per artifact type, and the explicit authority to send something back without it being an incident | Because the constraint moved here the moment drafting became free. Every rung below this one increases the volume of material arriving at a review process that is already the limiting factor |
| 2 | Research access and provenance discipline | A standing route to real customers that does not need a favour, plus the habit of labelling every claim measured, estimated by someone who would know, or invented, at intake rather than at review | Because review is only possible if claims carry their sources. Unlabelled evidence makes a reviewer do the research themselves, which converts your new review capacity back into production capacity |
| 3 | Instrumentation | An analyst, and the event work, so that a baseline and its variance exist before a build starts rather than being reconstructed afterwards | Because a metric you did not instrument is a metric you will not have at the readout. Cadence's 11% revisit rate with no target ever set is what an unfunded rung 3 looks like from the inside |
| 4 | Tooling | Seats, a workspace whose retention and training settings you can defend to security, and a place to put the material that must not go anywhere else | Because it is genuinely useful, genuinely cheap, and genuinely not a differentiator. It is also the only rung anyone will proactively ask you for, which is why it needs the discipline of going last |
LiveWhy review capacity is the first rung and not the fourth4 min▶
Run the arithmetic on your own team. If drafting a spec used to cost a PM two days and now costs an afternoon, you have not made the team four times faster. You have made the first draft four times cheaper while leaving the read, the challenge and the sign-off exactly as expensive as they were. Whatever your review throughput was last year, it is your throughput this year, and there is now four times as much arriving at it.
- The visible symptom is approval by recognition. A reviewer reads the shape of a document, recognises the shape, and approves it. That was survivable when producing the shape required making the decisions. It is not survivable now that the shape is free.
- The measurable symptom is a send-back rate near zero. Count it. If fewer than one document in five comes back for a named reason, review is being simulated rather than performed, and everyone involved is behaving reasonably given the time they have.
- The fix is boring and it is a budget line. Named reviewers, protected time in the week, and a norm that sending something back is the job rather than an escalation. It is headcount or it is hours, and it will not appear from goodwill.
LiveRung 3 is the one that gets cut, every time3 min▶
Instrumentation loses budget arguments for a structural reason rather than a bad one. It produces nothing a customer can see, it delivers value three months after it is paid for, and the cost of skipping it is invisible until the exact moment you need the number and it does not exist. Every organisation therefore discovers rung 3 at a readout, which is the most expensive possible place to discover it.
The leader move is to fund it out of the build rather than alongside it. When a bet is approved, the baseline pull and the event work are part of that bet's cost, not a separate request that has to compete with features. That single accounting change is worth more than any amount of advocacy, because it stops instrumentation being a thing somebody has to be brave about asking for.
The practitioner version of this rung is the PM track's b7, which writes the event spec for a real feature before any engineering starts. It is a useful thing to have read before you next approve a build, because it shows you exactly what you are declining to pay for.
Self-studyWhat rung 4 is actually worth, stated honestly3 min read▶
Nothing on this page argues that tooling does not matter. It matters, it is cheap relative to salaries, and a governed workspace with defensible retention settings is a real requirement rather than a nice-to-have. Buy it. The argument is only about sequence and about expectations.
Two things follow from tooling being last. First, the return on it is bounded by the three rungs above: better drafting into an organisation that cannot review, cannot source and cannot measure produces more confident wrong decisions, faster. Second, because every competitor can buy the same capability within a week, none of your advantage will ever live here. Your advantage lives in whether your organisation can tell a good decision from a well-written one, and that is rungs 1 to 3.
The uncomfortable corollary for a budget conversation: if the only line you can get approved this year is rung 4, you have not been given an AI budget. You have been given a software purchase, and you should say so out loud rather than let it be counted as a transformation.
Fewer, better decisions - not more documents 7 min live
If you fund the four rungs and change nothing else, the money is wasted, because your PMs are still being measured on the thing that just became free. A PM whose review counts specs written, tickets authored and turnaround time on document requests will respond to cheaper drafting in the only rational way available: by producing more. You will get four times the volume, the same number of decisions, and a review process quietly collapsing under material nobody asked for.
LiveThe five things to take off the review template this quarter4 min▶
Removing things is the harder half and the half that works. Each of these was a reasonable proxy when writing was the expensive part, and each is now actively harmful:
- Documents shipped. Once a competent draft is an afternoon, this measures typing speed. Worse, it rewards the PM who never sends anything back, because sending back is somebody else's document count.
- Responsiveness on document requests. The fast answer to "can you write me a one-pager on this" is now always yes, so the metric has stopped discriminating between people entirely.
- Coverage. "Every feature has a PRD" was a reasonable hygiene goal and is now trivially satisfiable with documents that decide nothing. It is the single easiest metric in this list to hit dishonestly without anyone intending to.
- Roadmap items delivered. Counts things shipped, not bets that paid. It quietly penalises the PM who killed their own feature at week three on the pre-agreed rule, which is behaviour you spent all of a5 trying to buy.
- Anything with the word volume, throughput or velocity attached to a PM rather than to a team. Delivery throughput is a real thing and it belongs to the delivery function. It has never been a measure of product judgement.
LiveWhat a good replacement measure looks like3 min▶
The right-hand column of the figure is not a scorecard, and it should not become one with weights and a total. Four of the five entries are counted in single digits per quarter, which is the point: they are small enough to discuss individually in a one-to-one rather than rolled into a number.
- Decisions with a written trade-off. Not decisions made, which is unfalsifiable. Decisions where the document names what was given up and who loses by it. Three per quarter from a senior PM is healthy; zero means they are ranking rather than choosing.
- Bets with a baseline that existed before the build. Binary, easy to check, and directly funded by rung 3. This is the single best proxy for whether the whole budget is working.
- Readouts that ended in a stop. Over a year, not a quarter. Zero over a year is the finding from a5, and it is about your culture rather than their competence.
- Specs signable at first review. The one measure that should improve over time and that maps onto a real cost you are paying. It rewards the PM who made the decisions before writing rather than the one who wrote first and negotiated after.
Notice that all four are checkable by someone other than the PM, in under five minutes, from artifacts that already exist. A measure that requires a new report to produce will not survive two quarters.
Self-studyWhat this does to the shape of the team3 min read▶
Two structural effects worth planning for rather than discovering.
First, the ratio of PMs to output changes without the headcount changing. A team of six that was producing enough documentation for six can now produce enough for twenty, so the question stops being "do we need another PM" and becomes "who reads what the six produce". That is a different hire, and often a more senior one. Resist the version of this conversation that concludes you need fewer PMs, because you have not made deciding cheaper, only writing.
Second, the junior role gets harder to define and easier to get wrong. The tasks that used to be a junior PM's apprenticeship - the first draft, the competitive summary, the notes writeup - are exactly the tasks that are now cheap. If you hand all of them to a tool, you have removed the ladder while keeping the job title, and in three years you will have a shortage of people who can tell a real spec from a fluent one. Part 3 is about that risk specifically, and it is the part of this session with the longest payback.
The craft risk, stated without the euphemism 7 min live
Here is the thing that does not appear in strategy decks about AI-assisted teams, and it should. A PM who has never had to write a spec from nothing has never met the blank page, and the blank page is where you learn what a decision costs to make. Without that experience they can read a fluent, well-structured, decision-free document and feel that it is fine, because they have no internal model of the resistance that making those decisions would have produced. The absence has no shape to them.
LiveWhy this is a real risk and not nostalgia4 min▶
The nostalgic version of this argument is that people should struggle because struggling builds character, and it deserves to be dismissed. The real version is narrower and testable:
- Recognising an absence requires having produced the thing that is absent. A PM who has written non-goals knows that writing them is where the argument happens, so a spec with no non-goals reads as unfinished to them. A PM who has only ever reviewed drafted specs sees a document with all its headings filled in.
- The failure is silent and slow. Nobody ever discovers it in a review. It surfaces one or two quarters later as a pattern of builds that went sideways for reasons that were all visible in the spec, and by then it looks like bad luck distributed across several people.
- It compounds through the review chain. If the reviewer cannot see the missing decision either, then rung 1 of your budget has bought hours of attention from someone who does not know what to look for, which is the most expensive way to buy nothing.
The containment is not to ban drafting help. It is to make sure judgement is exercised somewhere in every loop: the PM reviews explicitly against a named checklist rather than by feel, drafts are marked as drafts until a person signs them, and every PM writes some real specs unaided so the muscle exists. Two a quarter is enough. The point is not the two documents; it is that they cannot be produced without meeting the decisions.
LiveHow to interview for the shifted job4 min▶
The classic PM take-home is "write us a PRD for X". That exercise now measures almost nothing, because the candidate can produce a well-structured, confident, complete-looking document in the time it takes to read the brief. You are no longer testing what you think you are testing, and the strongest candidates know it, which makes the exercise slightly insulting as well as uninformative.
Replace it with a critique. Hand them a two-page spec that is fluent, well-formatted, plausible, and contains no decisions: a problem asserted with no evidence, "users" as the segment, "improve engagement" as success, no non-goals, happy path only, everything winning. Give them ten minutes and one question: what did this document decide, and what did it leave to somebody else?
- Strong signal: they find the missing non-goals within a minute and say so plainly. They notice that success is directional rather than measurable and ask what the number is today. They ask what was given up, and who loses. They ask who ends up deciding the edge cases, and answer their own question with "an engineer, on a Friday".
- Weak signal: they praise the structure and suggest adding a section. They rewrite it longer. They critique the prose, the formatting, or the tone. They ask for more context before they will comment, which sounds diligent and is usually an inability to see anything wrong with what is in front of them.
- Excellent signal, and rare: they ask you where the document came from and what you were hoping they would say. That candidate is reading the exercise as well as the artifact, and they will do that to your roadmap too.
The same exercise works as a development tool on people you already employ, and it is less confronting than it sounds because the document is obviously not theirs. Run it once a quarter with a different fake spec and you will learn more about your team's judgement than any competency framework will tell you.
Self-studyWhat this changes about hiring juniors3 min read▶
The temptation is to stop hiring junior PMs on the grounds that their traditional work is now free. It is a coherent short-term argument and a bad three-year one, because seniors are made out of juniors and nobody else is making them either.
The workable version is to change what a junior does rather than whether one exists. Their apprenticeship stops being "produce the first draft" and becomes "verify the draft against the source material" - reading the actual interview transcripts behind a synthesis, checking a claimed number against the dashboard, finding the check a spec fails. That work is genuinely valuable, it is exactly rung 1 and rung 2 of the budget, and it teaches judgement faster than drafting ever did, because it consists entirely of looking for the gap between what a document says and what is true.
Then give them the two unaided specs a quarter so they meet the blank page on purpose. A junior who spends a year verifying other people's work and writing eight specs from nothing is a better mid-level PM than the version of them that spent that year producing forty first drafts, and it is not close.
Rank your own spend, then find the missing line ◆ run this with your leadership team
This exercise takes twelve minutes and it usually ends with one uncomfortable discovery, which is that the first rung is not a line item anywhere in your budget. It is not a trick: review capacity is nearly always paid for out of goodwill and evenings, which is why it silently fails first.
Everyone writes down what the organisation currently spends on the four rungs, in order of size. Money, headcount or hours - any unit, as long as it is the same unit for all four. Two minutes, in silence, before anybody speaks.
Now rewrite the same list in the funding order from Part 1. The gap between the two orderings is the finding, and in most rooms it is close to a reversal.
Look specifically for review capacity as a line. If it is not there, it is not funded, and it is your bottleneck. Say the two numbers out loud: documents produced last month, documents sent back with a reason.
Pick exactly one rung to fund in the next thirty days, and name what loses that money. A funding decision with no loser is a wish, and everyone in the room already knows it.
Then say what you will stop measuring your PMs on, before you say what you will start measuring. Out loud, in this room, with the deletions first. If you cannot name three deletions, the new measures will land as additions and nothing will change.
LivePrompt 1 · the bottleneck question4 min▶
What a good answer sounds like: two actual numbers, and a send-back count that is not zero. "Roughly thirty documents, I read eleven of them properly, and I sent four back - two for a missing metric, one for no non-goals, one because the evidence was a Slack message." That answer contains a working review process and a person who knows what they are looking for. A very good answer volunteers the ratio without being asked and then tells you which artifact type never gets reviewed at all.
The failure mode to listen for: "everything gets reviewed", with no count. That is the answer that means review is happening as a calendar event rather than as an act, and it is almost never a lie - the meetings genuinely occur. Second tell: the send-back number is zero and nobody finds that surprising. Third and most diagnostic: they answer with how long they spend rather than what they found, because hours are the input and findings are the output, and only one of the two tells you the process is real.
LivePrompt 2 · what we stop measuring4 min▶
What a good answer sounds like: someone names a real one within a few seconds, and it stings slightly. "Coverage. Every feature having a PRD means I write PRDs for things that were decided in a corridor, and everyone knows the document is a formality." That is the answer you want, and the correct response is to remove the thing that week rather than to discuss it. Going first yourself is not a courtesy, it is what makes the second answer possible.
The failure mode to listen for: silence, then somebody offers a measure that is safely somebody else's. Second tell: the three things named are all inconveniences rather than distortions - process overhead, a tedious template - which means the room is answering a question about annoyance instead of about incentives. Third: everyone agrees enthusiastically and nothing is removed within the month, at which point you have taught your team that this kind of question is theatre, and the next one will be answered accordingly.
Self-studyPrompt 3 · the critique exercise, run on your own team4 min▶
What a good answer sounds like: the missing non-goals get named first, usually within a minute, followed by the observation that success is directional rather than measurable. Then the good ones go one step further and say who ends up making the decisions that are absent - the engineer who hits the edge case, the support lead who inherits the scope, the stakeholder who assumed the thing that is not excluded. That last move is the whole skill, because it converts a document review into a forecast of what will actually happen.
The failure mode to listen for: praise for the structure, followed by a suggestion to add a section. That is a person reading the shape and recognising the shape. Second tell: they rewrite it, at length, and the rewrite is also fluent and also decides nothing. Third: they ask for more context before they will comment. It sounds rigorous and it is usually the sound of someone who cannot see anything wrong with what is in front of them, because everything that should be there appears to be there.
The questions to ask your PMs ◐ 7 questions
Four of these belong in a document review and three belong in a one-to-one. None of them is a trap, and all of them are answerable in a sentence by somebody who did the work. Ask them routinely rather than when you are already suspicious, or they become an audit and the answers become performances.
- "Which of the decisions in this document did you make, and which did the draft make for you?" The single best question on this page. It is answerable in thirty seconds by anyone who wrote the thing properly, it is unanswerable by anyone who did not, and it never reads as an accusation because the premise is that using a draft is normal. A good answer names two or three decisions and points at them.
- "What did you send back last month, and what for?" Asked of a senior PM, this measures whether review is happening below you. A named artifact and a named reason is a working process. "Nothing needed sending back" is a finding about capacity, not about quality.
- "What is the baseline for this, and did it exist before the build started?" The rung 3 question in the room where it costs the least to fix. A metric with no current value cannot be missed, and a baseline pulled after the fact is a number chosen with the answer already known.
- "Where did this evidence come from - measured, estimated by someone who would know, or invented?" Three options, asked in a neutral tone, applied to every claim rather than the suspicious one. It normalises provenance as tidiness rather than as challenge, and the invented cells come out voluntarily once the label exists.
- "What did we decide not to do here, and who have you told?" Two questions in one, and the second half is the one that catches things. A non-goal nobody outside the team has heard is not a decision yet; it is an intention that will be relitigated by the first person who assumed otherwise.
- "When did you last write a spec from nothing, and how long did it take?" The craft question, asked warmly and infrequently. You are not checking productivity. You are finding out whether the muscle is still there, and the honest answer from a good PM is sometimes "not this year", which is information about how you have set the job up rather than about them.
- "If I doubled your tooling budget and halved your review time, would you be better off or worse off?" Ask this once and listen carefully. A PM who says worse off has understood the whole session and you should promote them into the reviewing role. A PM who says better off is telling you something true about where their bottleneck currently is, and it is worth finding out which of the four rungs they are actually short of.
Thirty days, and the first move costs nothing ◆ take this one away
Everything in this track has been judgement applied to somebody else's work. This is the part where you do something. The plan below fits in a month, needs no budget approval for the first two weeks, and is ordered so that the cheap diagnostic work happens before you spend anything.
30-DAY PLAN · the AI-assisted product org · owner: head of product
WEEK 1 MEASURE THE BOTTLENECK, SPEND NOTHING
Count last month's documents and how many were sent back with a
named reason. If the send-back rate is under 1 in 5, review is
being simulated rather than performed. Write both numbers down
where your leadership team can see them.
Also this week: name a reviewer for every artifact type. Not a
committee, a person, by name, per type. Cost: one hour.
WEEK 2 DELETE BEFORE YOU ADD
Take three volume measures off the PM review template. Say which
three, say that you put them there, and say it before you mention
any replacement. Suggested first cuts: documents shipped,
turnaround on document requests, and PRD coverage.
Then run the critique exercise once, on your own team, with a
fluent decision-free spec. Ten minutes each. Do not grade it.
WEEK 3 FUND ONE RUNG, AND NAME THE LOSER
Pick the highest unfunded rung, not the easiest one. In most orgs
that is rung 1 or rung 3. Attach an amount and say publicly what
is not getting that money, because a funding decision with no
loser will be quietly reversed by the first competing request.
If you fund rung 3, fund it INSIDE the next approved bet rather
than as a separate line, so it stops competing with features.
WEEK 4 MAKE IT STICK IN TWO SENTENCES
Add to the approval meeting, permanently: "what is the one number
and what is it today", and "what result would make us stop".
Both are from a5, both cost eight seconds, and both are unaskable
later. Then write one paragraph to the whole org explaining the
funding order and why tooling is last. Somebody will forward it
to a vendor. That is fine.
NOT IN Charters, milestones, WBS, critical path, risk registers, status
THIS PLAN cadence, rollout mechanics. All of that is the delivery function
and it has its own course. This plan stops at the decision.
The most common version of this budget going wrong does not look like waste, which is why it survives review. An organisation buys the tooling, runs the enablement, and gets a genuine and measurable increase in output: more competitive analyses, more one-pagers, more research summaries, more roadmap variants than anybody has produced before. The productivity slide is honest. The screenshots are real. Nobody is exaggerating anything.
What happens next takes about two quarters to become visible. The review meeting that used to cover four documents now covers eleven, in the same hour, so each one gets a skim and a nod. Two of the eleven contain a metric with no baseline and one contains a research theme that came from a single account, and none of the three gets caught, because catching them requires the kind of reading there is no longer time for. The builds that follow are not obviously wrong. They are ordinary, defensible, and slightly worse than the ones from before, and the retro on each of them finds a different local cause. Nobody ever writes the sentence "we increased our production capacity without increasing our review capacity", because no single quarter provides the evidence for it.
The tell is available on day one, and it is free: the ratio of documents produced to documents sent back. Watch that number rather than the output number. If production doubles and send-backs stay flat, your organisation has not got faster at deciding. It has got faster at agreeing.
Try it yourself - this week ◐ 30-40 min total
- Pull the two numbers: documents produced last month, documents sent back with a named reason. Do not share them yet. The ratio is your real review capacity and every other decision in this session scales off it.
- Write your current spend against the four rungs, then rewrite it in the funding order. Circle the rung that is missing entirely. In most organisations it is rung 1, and it is missing because nobody has ever had to ask for it.
- Take three volume measures off your PM review template this week, name them out loud, and say that you were the person who added them. Do the deletions before you announce a single replacement.
- Run the critique exercise once on your own team: a fluent, decision-free spec, ten minutes, one question. Note who finds the missing non-goals and who suggests adding a section. Do not grade it and do not tell anybody it was a test, because it is a development tool rather than an assessment.
- Fund exactly one rung in the next thirty days and name publicly what loses that money. Then write the paragraph explaining why tooling is last, and send it to the person most likely to disagree with you.
Sources covered
Full source map in materials/official-course-map.md. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · You have budget for exactly one of the four rungs this year. Which do you fund, and why that one?
Instrumentation and tooling are both real and both belong in the budget. But funding either one first raises the volume of material flowing into a review process that has not grown, which converts spend into more confident wrong decisions rather than better ones. The order is the argument.
2 · You fund all four rungs and leave the PM review template unchanged. What happens?
A PM measured on specs written, tickets authored and turnaround time will respond rationally to cheaper drafting by producing more of all three. The deletions from the review template are not a soft accompaniment to the budget; without them the budget funds volume, and volume is the thing you were already drowning in.
3 · You are hiring a senior PM. Which exercise tells you most about their judgement?
A good-looking PRD is now an afternoon's work at most, so the writing exercise no longer discriminates between candidates. Critique does: seeing an absence requires having produced the thing that is absent, which is exactly the craft at risk when drafting is always delegated.