learn-ai-pm-with-phoebe / PM session 2 of 10
Learn AI Product Management with Phoebe · PM track · Session 2 of 10

From a vague ask to an evidenced problem

Three people have asked Cadence for three different things, and all three asks are stated in different units - one in lost deals, one in ticket counts, one in vibes. You cannot choose between things measured in different units, so nothing about the decision is possible yet. This session turns each ask into a problem statement somebody could disagree with on the evidence, grades the evidence honestly, and leaves the CEO's request sitting on the page with an empty evidence line - which is a far kinder place for it than the meeting where nobody could say so out loud.

🟢 PM track PMs · founders · product leads Bring your carried backlog item 45 min
0-3 · Welcome 3-15 · Request, problem, solution 15-40 · Anatomy, evidence grades, the three asks rewritten 40-45 · Q&A
Part 0

Why the three asks cannot be compared yet

Session b1 put three asks on your desk. Sales wants live in-meeting notes, and offers "every demo asks for it" plus three lost deals. Support wants post-meeting summaries, and offers 214 tickets in three months tagged "cannot find what was decided" - the second most common tag in the queue. The CEO wants an agent, and offers nothing at all. Every one of those is a reasonable thing for that person to want, and not one of them is a problem statement.

The gap matters because ranking is impossible across different units. A lost deal is not commensurable with a ticket tag, and neither is commensurable with a strategic feeling. b6 does the prioritisation, but prioritisation is arithmetic on comparable inputs, and this session is where the inputs get made comparable. By the end, all three asks are written in the same six-part shape, and the differences between them are visible instead of rhetorical.

Live - presented in session Self-study - read after class ▶ Worked artifact - real text to judge Sources covered
★ What you walk out with today The six-part problem statement as a habit, three evidence grades you can name out loud without insulting anybody, and Cadence's three asks rewritten side by side - including the one whose evidence line is empty.
Part 1 · covers continuous-discovery interviewing practice

A request is not a problem 6 min live

Nobody brings you a problem. People bring you requests, and a request is the output of somebody else's private thinking: they noticed something, formed a theory about the cause, picked a fix, and handed you the fix. The problem is the middle step, and it never survives the journey. That is not a criticism of your stakeholders - it is what being helpful looks like. It does mean the first thing you receive is always the least useful of the three.

An ask arrives at the end of somebody's thinking, not the start of yours Request · what they said "Add live notes to the meeting" arrives as one sentence with a solution already inside it and their credibility attached Problem · what breaks who, doing what, what breaks how often, what it costs and how we know all of that someone could disagree with it Solution · one way to fix it one of several ways to do it cheap to argue about expensive to be wrong about and the only part they can see Asks arrive right to left. The work is walking them back to the middle column. Interrogate the request and you get an argument about the solution. Interrogate the problem and you get a conversation about evidence - the one conversation that can actually change somebody's mind, including yours.
🔍 Click to zoom - three different things, routinely treated as one

Written out for Cadence's three asks, the smuggled solution is easy to see once you look for it. The third column is deliberately hedged: at this point in the session it is a hypothesis about what the request might be pointing at, not a finding. Turning column three into something checkable is the rest of the page.

The ask, as it arrivedThe solution smuggled inside itThe problem it might be pointing at
Sales: "we need live notes in the meeting"a streaming transcription pipeline and an in-call UIsomething in the demo fails to land, and the prospect cannot picture the value while watching
Support: "we need post-meeting summaries"generated summaries, one per meeting, delivered somewhereafter a meeting ends, admins cannot locate the decision inside what was recorded
CEO: "we should have an agent"an autonomous system acting on meetings on a user's behalfunknown - no user, no moment, no failure and no frequency were named in the ask
LiveWhy every ask arrives pre-shaped as a solution4 min

There are three reasons, and knowing which one you are looking at changes how you respond:

  • A solution is concrete and a problem is abstract. "Add live notes" is a sentence anybody can picture. "Prospects cannot evaluate asynchronous value during a synchronous demo" is a sentence that gets you asked to speak plainly. People bring the version that communicates.
  • They have already done the thinking, just privately. The salesperson watched a prospect's face change. That observation is the most valuable thing in the room and it is not in the ask - it got compressed out on the way to you.
  • Naming a solution is how influence works in most companies. A request for a feature can be tracked, escalated and won. A description of a problem cannot. If your org only rewards feature requests, you will only receive feature requests, and that is an org design fact rather than a personal failing.
The one question that unpacks all three "What did you see that made you think of this?" It is not a challenge, it costs nothing socially, and it reliably returns the observation that got compressed out. Note that it asks about something that happened, not about what they believe would work.
Self-studyWhere AI genuinely helps at this stage3 min read

This is a drafting problem with a judgement problem hiding inside it, so the split from b1 applies cleanly. Three uses that are worth the time:

  • Draft the problem statement six ways. Same evidence, six framings, from narrow to broad. You are not looking for the best prose - you are looking for the framing whose implications you disagree with, because that disagreement is where your real belief about the problem is hiding.
  • Ask for the counter-argument. "Argue that this is not a problem worth solving this quarter, using only the numbers above." A tool with no stake in the outcome writes this far more honestly than you will, and it takes about a minute.
  • Ask what would have to be true. "List everything that must hold for this statement to be correct." You get back a list of assumptions, some of which you can check today and had not thought to check.

And the hard line: it must never supply the evidence. Ask why admins lose decisions and you will get five fluent, plausible reasons, none of which came from an admin. That failure has its own session, because it is the expensive one - b3 stages it deliberately.

Part 2 · the six-part shape

A problem statement someone could disagree with 7 min live

The bar is not "true". Almost anything vague is true. The bar is disagreeable: a colleague who knows the product should be able to read your statement and say "I think the frequency is wrong" or "that is not the segment where this bites". If the only available response is agreement, you have written something nobody can argue with because there is nothing in it to argue about.

1 · Who exactly Not: "our users" Team admins on paid workspaces of 5 to 50 seats. Not solo users. 2 · What they are doing Not: "using the product" Arriving at Monday's meeting knowing what was agreed before. 3 · What breaks Not: "it is hard to use" The decision is 3 lines inside a 40-minute transcript. 4 · How often Not: "sometimes" 6 of 8 admins interviewed, on every meeting. Not an edge case. 5 · What it costs Not: "frustration" A manual re-read, a retyped note, or the decision simply lost. 6 · How we know Not: "everyone says so" 214 tickets, 8 interviews, 11% revisit, 68% still take own notes. The test is not whether it sounds true. It is whether a colleague could disagree with it on the evidence. If the only available reply is "sure, sounds about right", you have written a mood in the shape of a problem.
🔍 Click to zoom - six parts, and the vague phrase each one replaces
LivePart 6 is the one that changes the room4 min

Parts 1 to 5 make the statement specific. Part 6 - how we know - is the one that changes what happens in the meeting, for a slightly uncomfortable reason: it is the part that can be empty, and an empty box on a shared page is a fact rather than an accusation.

  • It moves the argument off the person. "I do not think we should build an agent" is a position that has an opponent. "The evidence line for the agent is blank" is a description of a document, and the person who brought the ask can fill it in.
  • It makes the next step obvious. A blank evidence line is not a rejection - it is a research task with a named owner and a size. Sometimes it takes a day to fill.
  • It protects you later. When the quarter goes badly, the question is always "why did we build that". A written evidence line answers it, and so does a written blank one, in a different and more useful way.
Write the frequency before you write the cost A rare, expensive problem and a constant, cheap one look identical in prose and behave completely differently in a roadmap. Frequency is also usually the number you can get today from data you already have, which makes it the cheapest part of the statement to fill honestly.
LiveInterrogating an ask without making an enemy4 min

The whole technique is one substitution: ask for the instance, not the opinion. Opinions have to be defended, so asking about one starts a contest. Instances just get recounted, and recounting is a friendly act.

Do not askAsk insteadWhat you get back
"Do you really think that would have won those deals?""Which deals, and what did the buyer actually say?"three names and some quotes, or an honest "let me check"
"Is every demo really asking for it?""Can you send me the last five demos where it came up?"a countable set, which converts an assertion into data
"Why do you want an agent?""Who would use it, and what would they stop doing?"either a real user and moment, or a visible gap
"Do we have evidence for this?""What did you see that made you think of this?"the observation that got compressed out of the ask

Two habits make this land rather than sting. First, ask the same questions of your own favourite idea, out loud, in the same meeting - the questions read as a standard rather than a challenge once people have watched you fail them. Second, when someone cannot produce the instance, do not win. Write "to be checked" in the evidence line and offer to check it with them. You want the practice to survive the meeting, not the point.

Part 3 · covers the research-quality literature

The three grades of evidence a PM actually gets 6 min live

Product decisions almost never run on one clean kind of evidence. They run on a mixture of three, and the mixture is fine - what is not fine is a mixture presented as though it were all one grade. Naming the grade takes four words and removes most of the disagreement in a product review, because a surprising amount of that disagreement is two people arguing at different grades without noticing.

All three are legitimate. Only one of them is usually labelled honestly. Grade 1 · Counted analytics, ticket tags, logs 11% revisit rate. 214 tickets. Strong: it definitely happened Blind: it cannot tell you why Grade 2 · Reported interviews, surveys, diaries 6 of 8 never reopen a transcript Strong: it tells you why Blind: people misremember Grade 3 · Asserted "everyone is asking for it" "every demo asks for it" Strong: fast, and often correct Blind: no instance behind it Most escalations are grade 3 wearing grade 1's clothes. "Every demo asks for it" sounds counted. Ask which demos and you get three deal names, and nobody has yet checked what the buyer actually said. The fix is not to argue - it is to ask for the instance, not the opinion.
🔍 Click to zoom - three grades, and the disguise that causes most of the trouble
GradeWhat it isCadence exampleWhat it can and cannot settle
1 · counted behavioursomething a system recorded happening, countable without asking anyone214 tickets in 3 months; week-1 revisit rate of 11%; accuracy complaints on 1.4% of meetingssettles whether and how often. Cannot settle why, and cannot see anything you never instrumented
2 · reported behaviourwhat people say they did, in their own words, about a specific past occasion6 of 8 interviewed admins never reopen a transcript; 68% of surveyed admins still take their own notessettles why and what it costs them. Cannot settle scale, and drifts toward the tidier story
3 · asserted beliefa confident claim with no instance attached, often from someone with good instincts"every demo asks for it"; "everyone is shipping agents"; three lost deals offered but not yet examinedsettles nothing on its own. Useful as a pointer to where to go looking, dangerous as a reason to build
LiveGrade 3 is not the enemy - unlabelled grade 3 is4 min

A senior person's instinct is genuinely informative. They have seen hundreds of customer conversations and their pattern-matching is often right, which is exactly why the unlabelled version is dangerous rather than merely wrong.

  • Labelled, it is a hypothesis with a sponsor. "My read, no data yet, is that live notes are what closes the demo" is a useful sentence. It says what it is, and it tells you what to go and count.
  • Unlabelled, it outranks everything else in the room. Confidence plus seniority reads as evidence, and 214 counted tickets quietly lose to a strongly stated feeling. Nobody in the room did anything wrong; the units were simply never compared.
  • Escalation launders the grade. An assertion that travels through three retellings arrives sounding like a finding. "Sales says every demo asks for it" becomes "demos are asking for it" becomes "there is demand for live notes". Each step drops a hedge.
Four words that do most of the work Put "evidence grade" in your problem-statement template and fill it with one of: counted, reported, asserted, or none yet. It is not bureaucracy - it is the smallest possible change that lets a room compare two claims without anyone having to be the person who doubts a colleague.
Self-studyTwo traps inside grade 1 and grade 23 min read

Grading is not a ladder where higher always wins. Both of the upper grades have failure modes that a confident number happily hides:

  • Counted behaviour only counts what you instrumented. Cadence's 11% revisit rate is a real number about a real event, and it says nothing about admins who solved the problem outside the product entirely - the ones who retype three lines into a chat channel and never come back. The most important behaviour in a product is often the behaviour that leaves no event.
  • A ticket tag is a taxonomy decision, not a fact. "Cannot find what was decided" being the second most common tag partly reflects how the support team labels things. Before that number carries a quarter, someone should read thirty of the tickets. That is a two-hour job and it has changed the meaning of a headline number more than once.
  • Reported behaviour drifts toward coherence. People are not lying; they are constructing a tidy narrative from a messy week. This is why story-based interviewing asks "walk me through the last time" rather than "what do you usually do", and why "would you use this?" is not evidence of anything at all.

Which is the honest summary of the whole grading exercise: it tells you what a claim can carry, not whether the claim is correct. b3 takes the grade 2 pile - eight interviews - and shows how it fails when a synthesis is trusted without verification.

Worked artifact · 8 min · everyone

The three asks, rewritten ★ real text to judge

This is the output of the session: an opportunity brief with one entry per ask, each in the same six-part shape and each carrying its evidence grade. Read all three, then notice what the layout does on its own - nobody had to say the awkward thing, because the awkward thing is now a blank field on a shared page.

Take the ask verbatim first. Write it in the requester's words, with their name on it, before you touch anything. This is a courtesy that pays for itself: they will read the brief, and the first thing they look for is whether you heard them.

Extract the smuggled solution and set it aside. Name it explicitly on a separate line. Not to reject it, but so that it stops being the thing under discussion. It becomes a candidate again in b4, alongside alternatives the ask did not mention.

Fill the six parts, leaving blanks blank. Who, what they are trying to do, what breaks, how often, what it costs, how we know. Resist the urge to make a thin line sound thicker - the whole value of the brief is that thin lines look thin.

Grade the evidence in one word. Counted, reported, asserted, or none yet. Where grades are mixed, list each source with its own grade rather than averaging them into a vague "strong signal".

Write the cheapest next check. For every blank and every grade 3, one line: what would settle this, and roughly how long it takes. A brief that ends in three cheap checks is a brief that survives the meeting.

Brief 1 · Support's ask · evidence grade: counted plus reported

Ask, verbatim: "We need post-meeting summaries." (Support lead)  ·  Solution set aside: generated summaries, one per meeting.

Problem statement. Team admins on paid workspaces of 5 to 50 seats are trying to arrive at their next meeting knowing what was agreed in the last one. What breaks is that the decision exists as roughly three lines inside a 40-minute transcript, and the product offers no way to get to those lines without re-reading. This happens on effectively every recurring meeting: 6 of the 8 admins interviewed never reopen a transcript once the meeting ends, and the week-1 revisit rate across all recorded meetings is 11%. It costs them a manual re-read, a retyped note somewhere else, or the decision itself, quietly lost. We know this from 214 support tickets in 3 months tagged "cannot find what was decided" (the second most common tag), 8 interviews with paid-workspace admins, the 11% revisit rate, and a survey in which 68% of admins say they still take their own notes during meetings.

Cheapest next check: read 30 of the 214 tickets and confirm they are not concentrated in a handful of accounts. Half a day.

Brief 2 · Sales' ask · evidence grade: asserted, with three unexamined instances

Ask, verbatim: "Every demo asks for it - we need live notes in the meeting." (Sales)  ·  Solution set aside: streaming transcription with an in-call view.

Problem statement. Prospects evaluating Cadence in a live demo are trying to judge whether it will change their meetings. What breaks is unclear: the claim is that seeing nothing happen during the call weakens the demo, but no instance of that has been examined. Frequency is stated as "every demo" and is not counted. Cost is offered as three lost deals. We know this from a sales assertion plus three deal names; we do not know which deals, what the buyer actually said, or whether live notes were a factor in any of them.

Cheapest next check: read the notes on those three deals and ask the account owner what the buyer said, in their words. Under two hours, and it either upgrades this brief to grade 2 or it does not.

Brief 3 · The CEO's ask · evidence grade: none offered

Ask, verbatim: "We should have an agent." (CEO)  ·  Solution set aside: an autonomous system acting on meetings on a user's behalf.

Problem statement. Who: not specified. What they are trying to do: not specified. What breaks today: not specified. How often: not specified. What it costs: not specified. How we know: no evidence offered.

Cheapest next check: one 20-minute conversation, asking a single question - what did you see that made you think of this? Strategic instinct from a founder is a real signal about the market, and it is grade 3 until somebody names a user and a moment. That is a research task, not a verdict.

Real world

The empty field does the work that no human in the room wanted to do. In the meeting, "there is no evidence for the agent" is a sentence with a target, and most people will simply not say it to a founder. The same content, rendered as a row where five fields read "not specified" and the sixth reads "no evidence offered", is not an attack on anyone - it is a description of a document, sitting next to two other rows in the same format. Nobody has to be brave.

And notice what it does not say. It does not say do not build an agent. It says that as written, the ask cannot be ranked against a brief carrying 214 tickets and eight interviews, because there is nothing in it to rank. The next move is a research task with an owner and a size, which is a real answer and a much better outcome for the person who brought the ask than a polite yes followed by two quarters of nothing. The ranking itself is b6's job, and by then all three of these will be in the same units.

Practice

Your turn: grade, unpack, rewrite ★ 8 min

Three exercises on real-sounding material. Do the first two in session, out loud, with disagreement encouraged.

LiveQ1 · Grade these four claims3 min

One word each: counted, reported, asserted, or none yet. Say your answer before you read the verdict.

ClaimGradeWhy, and what it is worth
"Week-1 transcript revisit rate is 11% of recorded meetings."counteda real event count. Note there is no target and never has been, so 11% is a fact with no verdict attached. Whether 11% is bad is a judgement, not a measurement
"68% of surveyed admins still take their own notes during meetings."reporteda survey is self-report, so it carries why-ish information and unreliable scale. Strong here because it contradicts the product's core promise, which is a hard thing for a survey to produce by accident
"Enterprise buyers will not adopt this without live notes."assertedconfident, plausible, no instance. Worth exactly one question: which buyer said that, and when?
"Admins would prefer a summary in their chat tool."asserted, disguised as reportedthe trap. "Would prefer" is a prediction about the future, and no amount of asking makes it evidence. If one admin said it, that is one anecdote about one workflow. b3 takes this exact failure apart

The fourth is the one worth arguing about. "Would you use it" and "would you prefer" produce answers that sound like reported behaviour and are actually asserted belief - the interviewee's, not the stakeholder's, but asserted all the same.

LiveQ2 · Unpack an ask without the fight3 min

An engineering lead says: "we should let admins export everything to CSV, people keep asking." Write the three questions you would actually ask, then compare.

  • "Which people, and where did they ask?" - converts "keep asking" into a countable set, or reveals there is not one. Neutral, because it sounds like you are going to go and read them.
  • "What did the last person say they were going to do with the file?" - the instance question. Export is almost always a workaround for something the product does not do, and the workaround tells you the job. This is the question that most often changes what gets built.
  • "What made you think of this now?" - the compressed observation. Something happened this week.

Notice that none of the three contains the word "why", and none of them requires the other person to defend a position. Also notice what you did not do: you did not say no, and you did not say yes. You collected the instance, and the ask either gets stronger or quietly stops being urgent.

Self-studyQ3 · Your carried item, in six parts4 min

Take the backlog item you picked in b1 - the one that gets argued about and never decided - and write it in the six-part shape. Fill what you can, leave the rest genuinely blank, and grade each source in one word.

Then read the blanks, because they are the finding. In most cases the reason the argument keeps restarting is visible in exactly one place: either the frequency is unknown, so nobody can size it, or the whole thing rests on grade 3 that everybody has been treating as grade 1. If you can fill the blank in under a day, you have found the cheapest useful piece of work on your desk this week. If it would take a month, that is a real answer too, and it belongs in the brief.

Homework

Try it yourself - this week ◐ 20-30 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

Continuous-discovery interviewing practice - story-based questions, why "would you use it" is not evidencePart 1 + Part 2 · ask for the instance, not the opinion
The evidence standard - counted, reported and asserted claims, and what each can carryPart 3 · the three grades, and the grade-3 disguise
Opportunity-solution-tree thinking - the request as one solution among severalPart 1 · the smuggled solution; the tree itself is b4
Research-quality literature - sampling, instrumentation blind spots, coherence driftPart 3 self-study · the full synthesis treatment is b3
PRD practice - problem evidenced, the first of the six spec checksNamed here, taught and scored in b5
Delivery: charters, milestones, critical path, status reportingOut of scope by design - learn-ai-project-management
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · A stakeholder says "we need live notes in the meeting". What have you actually received?

Requests are the end of somebody else's thinking. The valuable part - what they saw that made them think of this - got dropped, because a solution is concrete and communicates, and a problem is abstract and does not. Your first move is to go and get the observation back.

2 · "Every demo asks for it, and we lost three deals over it." What grade is this, and what is the next move?

Most escalations are grade 3 wearing grade 1's clothes. A number inside an assertion does not make it counted; "every" is uncounted and the three deals have not been read. The whole move costs under two hours and it is a favour to the person who brought the ask, not a challenge to them.

3 · In the opportunity brief, the CEO's agent has "no evidence offered" written in its evidence line. What has that achieved?

A blank evidence line is not a verdict. Founder instinct is often right and is genuinely informative about the market - it is simply grade 3 until somebody names a user and a moment. Written down beside two other briefs in the same format, it becomes comparable and cheaply checkable, and nobody in the room had to be the one to doubt a colleague out loud.

PM session 2 cheat sheet · pin this

Three different thingsRequest (what they said) · problem (what breaks) · solution (one way to fix it). Asks arrive as the third.
The six partsWho exactly · what they are trying to do · what breaks · how often · what it costs · how we know.
The barNot "true" - disagreeable. If the only reply is "sounds right", you wrote a mood.
Three gradesCounted (analytics, tickets) · reported (interviews, surveys) · asserted ("everyone is asking").
The disguiseMost escalations are grade 3 wearing grade 1's clothes. A number inside an assertion is still an assertion.
The one substitutionAsk for the instance, not the opinion: "which deal, and what did they say?"
AI hereSix framings, the counter-argument, what would have to be true. Never the evidence itself.
The empty fieldA blank evidence line is a research task with a size, not a rejection. Next: b3, and how synthesis invents.