learn-ai-pm-with-phoebe / PM session 5 of 10
Learn AI Product Management with Phoebe · PM track · Session 5 of 10

The spec, and the scorer

A drafting tool will write you a PRD from one sentence, and it will be well-structured, confident, and worth almost nothing - because a spec is not a document, it is a set of decisions that happen to be written down. This session names the six decisions a spec has to contain, then scores five real drafts against them: the same tool, the same product, five different amounts of input. The score goes 15, 35, 55, 70, 100 - and the interesting column is not the score, it is which specific failure each new input removed.

🟠 PM track PMs · founders · product leads Signature lab · scores real text 45 min
0-3 · Welcome 3-14 · The six decisions 14-40 · The lab: run the ladder, score your own 40-45 · Q&A
Part 0

Why specs are where AI is most dangerous

Sessions b2 to b4 did the work: an evidenced problem, verified research, a named job, a chosen branch of the opportunity tree. This session turns that into the artifact someone builds from - and it is the point in the process where AI helps most and hurts most, for the same reason. It writes a document that looks finished. A reviewer reads the shape, recognises the shape, and approves it. Nobody notices that four of the six decisions were never made, because the sections where those decisions should live are full of confident prose.

Live - presented in session Self-study - read after class ▶ Live lab - scores real text, including yours Sources covered
★ What you walk out with today The six checks as a habit, a measured answer to what each input actually buys, and a scored version of one of your own real specs - with the failing check named.
Part 1 · covers PRD practice, the spec-as-decision-record framing

The six decisions a spec must contain 7 min live

Not sections - decisions. A spec can have all the right headings and none of these, which is exactly what rung 1 in the lab below demonstrates. Each of the six has a failure mode that is specific, common, and expensive, and each one gets decided by somebody whether or not you decide it.

1 · Problem evidenced Missing: asserted, not evidenced Then: it cannot be ranked against anything, so loudness decides 2 · User named Missing: "users" as a segment Then: you cannot say who it is NOT for, so scope has no edge 3 · Success measurable Missing: "improve engagement" Then: nobody can be wrong later, which sounds safe and is not 4 · Non-goals stated Missing: no not-doing list Then: everything adjacent is arguably in, one ticket at a time 5 · Edge cases decided Missing: happy path only Then: an engineer decides at 4pm Friday, and never tells you 6 · Trade-off explicit Missing: everything wins Then: it was never a decision, and the retro finds that out Every one of the six gets decided either way. The only question is whether you decided it, in writing, before the build - or whether it was decided later, quietly, by whoever hit it first.
🔍 Click to zoom - six decisions, six failure modes, and who decides when you do not
LiveThe two that are almost always missing4 min

In the lab below, non-goals and the trade-off survive four of the five rungs. That is not an accident of the example - it is the pattern, and both have the same cause: they are the two checks that require somebody to give something up.

  • Non-goals. A drafting tool has no reason to exclude anything, because exclusion is not a language problem. You have to hand it the scope decision, which means the scope decision has to exist first. If you cannot write the non-goals, you have not finished deciding - you have finished describing.
  • The trade-off. The hardest line in any spec is "we chose X at the cost of Y, and we are accepting that". AI will not write it unsolicited, because the trade-off is not in the input unless you put it there. A spec where everything wins is a spec where nothing was chosen, and it will read beautifully right up until the review where two people discover they agreed to different things.
The signable test Would an engineer who has never been in your meetings build the right thing from this, without asking you a question or inventing an answer? If no, it is not a spec yet - regardless of length, formatting, or how confident the prose sounds.
Self-studyWhat the six checks cannot tell you3 min read

The lab is a completeness check, not a quality check, and the difference matters:

  • It cannot tell you the metric is the right metric. A spec targeting a vanity number with a baseline and a target scores full marks on check 3. The scorer sees that a metric exists; it cannot see whether moving it means anything.
  • It cannot tell you the evidence is good evidence. Eight interviews from your friendliest customers pass check 1 as easily as eight representative ones. Session b3 is where that judgement lives.
  • It cannot tell you the trade-off was the right trade-off. It only sees that one was stated.
  • And it cannot tell you the problem is worth solving at all. A perfectly complete spec for the wrong problem scores 100.

Which is the whole reason this is a lab in the middle of a course rather than a tool you install. The six checks catch absence, reliably and cheaply, and absence is what an AI-drafted spec suffers from. Judgement is still yours, and it is the part nobody can score.

What a perfect score cannot tell you THE RIGHT METRIC? Scores full marks on a vanity number GOOD EVIDENCE? 8 friendly interviews pass like 8 real ones THE RIGHT TRADE-OFF? Only sees that one was stated at all WORTH SOLVING? A complete spec for the wrong problem: 100 The six checks catch absence, reliably and cheaply - judgement is still yours to make.
🔍 Click to zoom - a perfect score and a wrong answer are not mutually exclusive
Signature lab · 20 min · everyone

Run the ladder, then score your own ★ the scoring is real

Five rungs. Each one is the spec a drafting tool produces from that much input about Cadence's summaries feature and no more. Read the failing checks at each rung before you move up, and read the gained column when you run the whole ladder - that column is the actual lesson.

Start at rung 1 - one sentence in. Read the spec: it has every section a PRD should have, and 15 out of 100. Five of the six decisions were never made.

Add research (rung 2). The problem becomes checkable and the user gets named - two checks, from one input. Success is still a mood.

Add the metric definitions (rung 3), then constraints (rung 4). Watch success become measurable, then watch the edges get decided. Non-goals and the trade-off are still missing at 70 out of 100.

Add the non-goals (rung 5). 100 out of 100, signable - and note that one input fixed two checks, because writing down what you are not doing is what forces the trade-off into the open.

Three points on the ladder: 15, 70, 100 Rung 1 - one sentence 15 Rung 4 - + constraints 70 Rung 5 - + non-goals 100 One rung, non-goals, fixes two checks at once - it forces the trade-off into the open.
🔍 Click to zoom - the last rung is worth thirty points because it forces a trade-off, not because it is longer

Then press "Score your own spec" and paste something real. This is the part that matters after today.

Real world

The rung-5 spec is not longer because it is padded. It is longer because it contains four decisions the rung-1 version left to somebody else: who this is not for, what happens when a meeting has no clear decisions, whether summaries survive the 30-day retention setting, and the fact that this ships without the live-notes feature sales asked for. Those four sentences are the spec. Everything else is context around them - and a drafting tool will produce the context all day without ever producing the four sentences, because they are not writing, they are deciding.

Practice

Your turn: three specs, three verdicts ★ 8 min

Use the "Score your own spec" box in the lab above for each of these. Predict the failing check before you press Score.

LiveQ1 · The one-liner everyone has written3 min

Paste this in and score it:

Paste this PRD: dark mode. Problem: users want dark mode, it would improve satisfaction. Solution: add a toggle in settings.

It scores 0 out of 100 - all six checks fail. Worth sitting with for a second, because most teams have shipped from something very close to this, and the reason it scores zero is not that it is short. It is that there is nothing in it anyone could disagree with on the evidence, no way to know afterwards whether it worked, and no boundary at all.

LiveQ2 · Your own worst spec4 min

Find a spec you wrote when you were busy - ideally one that shipped and went sideways. Score it. Then look only at the failing checks and ask which of them explains what actually went wrong.

In practice it is usually one of two: the trade-off was missing, so two stakeholders agreed to different things; or success was not measurable, so the argument afterwards was about whether it worked rather than what to do next. Both are visible in the spec before a line of code exists, which is the uncomfortable part.

Self-studyQ3 · Score a spec you did not write3 min

No paste needed if you would rather not - a thought exercise. Take the last spec somebody handed you and run the six checks by eye. Then, and this is the actual skill, work out how to say the failing one out loud without it landing as criticism.

"What are we not doing here?" and "how will we know this worked?" are the two most useful questions in product review, and neither reads as an attack. Both are just checks 4 and 3, asked politely. A team that gets asked those two questions every time starts writing them into the spec before the review, which is the entire goal.

Homework

Try it yourself - this week ◐ 25-35 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

PRD practice - the spec as a decision record rather than a documentPart 1 + the lab · the six checks, measured across five drafts
Scope control through explicit non-goalsPart 1 · why non-goals survive four rungs, and what fixes them
Metric definition and instrumentationCheck 3 only · the full treatment is session b7
Research quality and representativenessCheck 1 sees that evidence exists, not whether it is good - that is b3
Estimation, sequencing and delivery planningOut of scope by design - learn-ai-project-management
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · Rung 1 of the lab scores 15 out of 100 but has a problem section, a solution section, a success section and requirements. What is missing?

A spec is a set of decisions that happen to be written down. Sections are where decisions live, not the decisions themselves - and a reviewer who reads the shape and recognises the shape will approve a document containing none of them.

2 · Adding the non-goals at rung 5 fixed two checks at once. Which, and why?

The two checks that survive longest are the two that require somebody to give something up. Once the not-doing list exists, the trade-off is one sentence away; without it, no amount of drafting produces one.

3 · A spec scores 100 out of 100. What does that guarantee?

The scorer catches absence, cheaply and reliably, which is what an AI-drafted spec suffers from. Whether the decisions were good ones is judgement, it is yours, and nothing on this page can score it.

PM session 5 cheat sheet · pin this

A spec isA set of decisions that happen to be written down. Sections are where they live, not what they are.
The sixProblem evidenced · user named · success measurable · non-goals stated · edges decided · trade-off explicit.
The ladder15 → 35 → 55 → 70 → 100. Read the gained column, not the score.
What each input buysResearch: evidence + user. Metrics: measurable. Constraints: edges. Non-goals: scope + trade-off.
The last twoNon-goals and the trade-off survive longest, because both require giving something up.
Signable testCould an engineer who was in none of your meetings build the right thing without inventing an answer?
Two polite questions"What are we not doing?" and "how will we know it worked?" - checks 4 and 3, asked kindly.
What 100 does not meanComplete is not correct. A perfect spec for the wrong problem scores 100. Next: b6, prioritisation.