learn-ai-pm-with-phoebe / PM session 3 of 10
Learn AI Product Management with Phoebe · PM track · Session 3 of 10

Research synthesis, and the trap

This is the highest-risk thing a PM does with AI, and it does not feel risky at all - which is the whole problem. Paste in eight interviews, get back six clean themes in seconds, each phrased more confidently than anything you would have written yourself. On this page, two of those six themes do not survive verification: one has zero sources behind it and one is a single person's sentence rewritten in the plural. The material is on the page, the themes are on the page, and the reveal is behind a card. Find them first.

🟡 PM track PMs · founders · product leads Bring your carried backlog item 45 min
0-3 · Welcome 3-13 · What synthesis is for 13-40 · The flow, the trap, the verification discipline 40-45 · Q&A
Part 0

Why this is the dangerous one

b2 left blanks in the evidence lines on purpose, and blanks are uncomfortable. Synthesis is where the temptation to fill them lives, because a tool that writes fluent themes will fill any blank you point it at, and it will never once tell you which of its themes came from your material and which came from its sense of how products like yours usually work. Both arrive in the same font, in the same confident register, in the same numbered list.

Everything else in this course has a cheap failure mode. A bad spec gets caught in review. A wrong priority gets caught in a quarter. An invented theme does not get caught, because it becomes the reason for the roadmap, and by the time the feature underperforms the theme has been repeated in four decks and nobody remembers that no customer ever said it. That is why this session exists and why it is the one to run live with your team rather than read alone.

Live - presented in session Self-study - read after class ▶ Worked artifact - real snippets and real themes Sources covered
★ What you walk out with today A four-step synthesis flow where only the last step matters, the two shapes of synthesis failure by name, and Cadence's six themes reduced to four verified ones plus one labelled anecdote plus one that never existed.
Part 1 · covers continuous-discovery interviewing practice

Synthesis reduces the number of things 6 min live

Most people asked to "synthesise the research" produce a summary of each interview. Eight interviews in, eight tidy paragraphs out. Every paragraph is accurate and the exercise has achieved nothing, because a pattern is not a property of an interview - it is a property of the set. It only appears when you stop looking at accounts one at a time.

Eight summaries is eight times the reading and none of the answer Summarising · one output per input 8 tidy paragraphs, one per admin interview every one of them accurate and quotable no pattern anywhere, because a pattern is a property of the set, not of any one account Output: a folder nobody reopens Synthesising · fewer outputs than inputs 4 themes plus 1 labelled anecdote each carrying a source count and a verbatim each traceable back to who actually said it and the weak one visibly marked as weak Output: something b4 can build a tree from The test is arithmetic: if your output has as many items as your input, you summarised. Eight interviews should not produce eight findings. They should produce three to five, one of which is probably an anecdote, and the count beside each one is what makes the next four sessions possible.
🔍 Click to zoom - the difference is a count, not a writing style
LiveWhat a theme has to be, to be worth the name4 min

A theme is a claim that a pattern exists across independent people. That definition has three consequences that most synthesis output quietly ignores:

  • It needs a count. "Admins struggle to find decisions" is not a theme until it reads "6 of 8". The count is not decoration - it is the entire difference between a finding and a phrase, and it is what makes two themes comparable later.
  • It needs independence. Three quotes from the same account is one source. Two admins who work at the same company and sat in the same meeting are closer to one source than two. This is where small research sets fail most often and least visibly.
  • It needs to be falsifiable by the material. If you cannot say which snippet would have to be absent for the theme to collapse, the theme is not attached to the material at all. It is attached to your prior beliefs, which are not research.
The shape a theme should be written in One sentence, then a slash, then the count, then one verbatim. "Admins do not reopen transcripts after a meeting ends / 6 of 8 / 'once the call is over I never open the transcript again'." Everything in that line is checkable in under a minute, which is the standard from b1 doing real work.
Self-studyWhere AI is actually useful in synthesis3 min read

The honest answer is that it is useful in three of the four steps, and it is very good at them:

  • Coding. Tagging thirty snippets with what each one is about is mechanical, tedious, and exactly the kind of work where a tool beats a tired human at 6pm. Check the codes, not the themes - codes are cheap to verify because each one points at one snippet.
  • Grouping. Proposing which codes cluster is a genuinely useful second opinion, and it will occasionally group two things you had mentally kept apart for no reason other than that you interviewed them on different days.
  • Arguing against your own theme. "Here are my four themes and the snippets. Which theme is least supported by this material?" This is the single most valuable prompt on this page, and it works because it asks for a critique of a fixed set rather than the generation of a new one.

What it cannot do is the fourth step. Verification is not a language operation - it is the act of a named person asserting that a pattern is real, and that assertion is what everything downstream rests on. A tool cannot make it, because a tool cannot be accountable when the pattern turns out not to exist.

Part 2 · the flow, and the step everyone skips

De-identify, code, theme, verify 6 min live

Four steps, in this order, every time. Three of them are handling and pattern-matching and can be delegated almost entirely. The fourth is a claim, and a claim needs an owner. The order matters as much as the steps: coding before theming stops you from reading the material through a theme you already had in mind, which is the most common way a synthesis confirms what the PM already believed.

Three steps you can hand over, and one you cannot 1 · De-identify names, employers, contacts out once, at intake Protects: consent, and you Skipped: never, legally 2 · Code tag what each snippet is actually about Protects: the raw material Skipped: often, quietly 3 · Theme group codes that recur across different people Protects: nothing yet Skipped: rarely, it is easy 4 · Verify each theme back to 2 or more sources, with a quote Protects: everything after Skipped: almost always Step 4 is the only one that cannot be delegated, and it is the only one that gets skipped. Steps 1 to 3 are handling and pattern-matching. Step 4 is the assertion that a pattern is real, made by a named person who will still be here in two quarters when the feature built on it underperforms.
🔍 Click to zoom - the order is load-bearing, and so is the last box
StepWhat you actually doDelegate?What goes wrong if you skip it
1 · de-identifystrip names, employers and contact details at intake, once, before anything else touches the materialthe mechanics, yes; the decision to do it, noyou have put a customer's identity into a tool your research consent almost certainly never covered - see b1's never-paste list
2 · codetag every snippet with what it is about, in the interviewee's terms rather than your feature namesyes, then spot-check about a fifth of themyou skip to themes and read the material through the theme you arrived with, which is confirmation with extra steps
3 · themegroup codes that recur across different people, and name each group in one plain sentenceyes, and a second opinion here is genuinely usefullittle, on its own - this is the step that feels like the work and carries the least risk
4 · verifytrace every theme back to at least two independent sources and paste one verbatim beside itnevereverything. This is the step where the invented theme and the over-generalised theme are caught, and it is the step that gets cut when the readout is tomorrow
LiveDe-identify at intake, not per prompt3 min

b1 put this on the never-paste list; here is the operational version. The rule is that de-identification happens once, when the transcript arrives, and produces the only copy that anybody works from afterwards.

  • Replace, do not delete. "Admin 3, 40-seat agency, marketing" carries every bit of product signal you need and none of the identity. Deleting instead of replacing loses the segment information, and the segment is the thing b2 spent a whole session insisting on.
  • Do it at intake because habits fail under deadline. A step you have to remember before each prompt is a step you will skip in the week the readout is due, which is exactly the week you are pasting the most material.
  • Keep the mapping somewhere else. You will need to go back to a specific admin for a follow-up. A separate key, held where your company holds customer data, gives you that without the identity ever travelling with the quotes.

This is also the step that makes the rest of the session legitimate at all. Everything below - eight admins, coded, themed, verified - is material that never needed a name.

Self-studyWhy coding before theming is not bureaucracy3 min read

The temptation is to go straight from transcripts to themes, and it is a strong one because themes are the deliverable and codes are not. The reason to resist is specific: a code is a claim about one snippet, and a theme is a claim about the set. Claims about one snippet are cheap to check, and claims about the set are not.

  • Codes are individually verifiable in seconds. Read the snippet, read the code, agree or disagree. You can check thirty codes in ten minutes, which is why delegating the coding is safe in a way that delegating the theming is not.
  • Codes stay in the interviewee's language. "Asks a colleague instead of scrolling" is a code. "Retrieval friction" is a theme wearing a code's clothes, and once it is in your notes as a code you will never see the raw behaviour again.
  • Codes make the count honest. When themes come first, the count is assembled afterwards to support the theme, and that is a different and much weaker activity than counting codes that already exist.

Worth being honest that this is slower. Coding eight interviews properly is an afternoon. Going straight to themes is twenty minutes. The afternoon is the price of a synthesis that survives someone pulling on it in a review.

Part 3 · covers the research-quality literature

Two shapes of failure, one detector 5 min live

Synthesis failures are not random. They come in two shapes, and both are shapes you can look for on purpose. Neither one looks like an error on the page: the invented theme reads as the most insightful item in the list, and the over-generalised theme reads as the most actionable. That is not a coincidence - a theme with nothing holding it down is free to be as clean as the language allows.

Two shapes of failure, and the single column that catches both Shape A · the invented theme Sources: zero Nothing in the material says it. It is simply what this kind of product usually has trouble with, so it arrives sounding like the most insightful item Detector: ask for the verbatim. There is not one Shape B · the over-generalised theme Sources: one One person's sentence, rewritten in the plural as "admins want ...". Real, quotable, and a whole strategy hanging off a single workflow Detector: ask for the second source. There is none One column catches both: the source count, printed beside every theme before anyone reads it. Zero sources is fiction. One source is an anecdote and stays labelled as one. Two independent sources is the floor for anything called a theme - and the PM reads the two quotes, not the sentence claiming they exist.
🔍 Click to zoom - both failures are invisible in prose and obvious in a count
LiveThe verification discipline, in four rules4 min

Short enough to remember, and the whole point is that they are mechanical rather than clever:

  • Every theme carries its source count. Written as "6 of 8", not "most" and not "many". The count goes in the same line as the theme, in every document the theme ever appears in, including the deck.
  • Every theme carries at least one verbatim. A real sentence a real person said, in their words. If nobody can produce one, the theme does not exist, and this catches shape A on its own.
  • A theme with one source is labelled an anecdote and never promoted. Not deleted - anecdotes are useful, and single-source signals are often the earliest sight of something real. Labelled. It can be promoted later by more research, never by being restated more confidently.
  • The PM reads the quotes, not the summary. This is the one that costs time and the one that cannot be substituted. Reading the quotes is how you notice that two of them are the same person, or that the sentence supporting your best theme is actually about something else.
The two questions that do the work in a review "How many people said this?" and "can you show me one of them saying it?" Neither is aggressive, both are answerable in seconds when the discipline was followed, and both are unanswerable when it was not. A team asked these twice starts putting the count in the document before the review, which is the actual goal.
Self-studyWhy the invented theme is always plausible3 min read

It helps to understand where an invented theme comes from, because it stops you from expecting it to look wrong. A drafting tool has seen an enormous amount of writing about products of this general kind: what their users complain about, what their reviews say, what their competitors fix. Asked to produce themes from your material, it produces themes that are consistent with your material and with everything it already knows about products like yours - and it has no mechanism to tell you which of the two produced any given line.

  • It will be a real problem, somewhere. The invented theme is usually a genuine issue for this product category. It is simply not one your eight people raised, which makes it a fine hypothesis and a terrible finding.
  • It will fit the seam between two real quotes. Invented themes tend to appear exactly where your material is thin, and thin patches attract plausible filler the way a gap in a hedge attracts a shortcut.
  • It will often be the one you like best. Unconstrained by evidence, it can be phrased perfectly. The cleanest sentence in a theme list deserves the most suspicion, which is an uncomfortable heuristic and a reliable one.

And be honest about the cost of catching it. Verification is slower than not verifying, every single time, with no compensating speed-up anywhere. That is the trade you are making: an afternoon now, against a quarter spent on a theme nobody said.

Worked artifact · 12 min · everyone

Cadence's eight interviews, and six themes ★ two do not survive

Below is the de-identified material from Cadence's 8 admin interviews with the codes already assigned, followed by the six themes an assistant produced from exactly this material and nothing else. Two of the six do not survive verification. One is invented and one is over-generalised from a single quote. Work out which two before you open the reveal in the practice section - and notice that the theme list, as produced, does not carry a single source count.

SourceVerbatim, de-identified at intakeCodes assigned
A1 · 12-seat studio"Once the call is over I never open the transcript again. I write three lines into our project channel while it is still fresh."no-reopen · writes-own-summary-elsewhere
A2 · 40-seat agency"The thing I need is what we agreed. It is in there somewhere around minute 34, but I have to go and find it."decision-is-the-unit · hard-to-locate
A3 · 18-seat consultancy"I still take my own notes. I know that is the whole point of the tool. I just do not trust that I will be able to find it later."parallel-note · low-trust-in-retrieval
A4 · 25-seat product team"Monday standup is the problem. I need to arrive knowing what last week's decisions were, and I usually reconstruct that from memory."pre-meeting-need · reconstructs-from-memory
A5 · 9-seat agency"I have reopened a transcript maybe twice, both times because somebody disputed what was said."reopen-only-on-dispute
A6 · 50-seat services firm"Nobody on my team reads them. I share the link and it sits there."shared-not-read
A7 · 15-seat studio"It is faster to ask the person than to scroll back through forty minutes. So I ask."asks-a-human-instead · decision-is-the-unit
A8 · 30-seat agency"I would want it in the channel we already use. If I have to log in somewhere to read a summary, I will not."wants-it-where-they-work · login-friction
The six themes, exactly as produced from the material above

T1. Admins do not reopen a transcript once the meeting has ended.

T2. The unit of value is the decision, not the recording.

T3. Admins keep a parallel manual note because they do not trust they will find things later.

T4. Admins want summaries delivered into the tools they already work in, rather than inside the product.

T5. Transcript accuracy is the main barrier to trust and adoption.

T6. Admins reconstruct last week's decisions from memory before recurring meetings.

Read that list on its own and it is genuinely good. Six clear sentences, no hedging, no repetition, in an order that almost tells a story. It took seconds, it is better written than most synthesis produced by hand at the end of a long week, and four of the six are correct. The problem is that nothing on the list distinguishes the four from the two, and there is no way to find out by reading it more carefully. You have to go back to the material.

Take the theme list and add an empty column. Head it "sources", and do not let yourself fill it from memory. This one mechanical act is what turns a list of sentences into something you can audit, and it is why the count belongs in the template rather than in your good intentions.

Go theme by theme, not source by source. For T1, walk all eight rows and write down every source ID that supports it. Then T2. Working the other way round - reading each snippet and deciding which theme it fits - is how you accidentally assign a snippet to a theme it merely rhymes with.

Paste one verbatim beside each theme. Not a paraphrase, and not your own sentence that captures the spirit of it. If you find yourself writing the quote rather than copying it, you have found a theme with no quote, which is the finding.

Check independence, not just count. Two supporting snippets from the same admin is one source. Two admins at the same 50-seat firm who sat in the same meeting is closer to one than two. Small research sets fail here more often than anywhere else.

Apply the floor and relabel. Zero sources: delete it and write down that it was produced, because that is useful information about your process. One source: relabel it as an anecdote with the source ID attached. Two or more independent sources: it is a theme, and it goes to b4 with its count.

Real world

The reason this failure survives is that the theme list gets separated from the material almost immediately. The themes go into a slide. The slide goes into a review. Somebody restates one of the themes in a planning doc, and by then it is a sentence with no count, no quote and no source, in a document whose author was three steps away from the transcripts. Nothing in that chain is anyone's mistake, and by the fourth retelling the theme is simply what the company believes about its users.

Which is why the count travels with the sentence or the discipline does not exist. "6 of 8" is four characters of overhead and it survives being copied into a deck, whereas "we heard consistently that" does not survive anything. When a theme arrives at a planning meeting without a count, the honest response is not to doubt the person presenting it. It is to ask the two questions - how many, and can you show me one - and to notice how quickly a room learns to bring the answer.

Practice

Your turn: find the two, then check your own ★ 8 min

Do Q1 with the material still on screen and commit to two theme numbers out loud before anybody opens Q2. Guessing in public is the point of the exercise.

LiveQ1 · Build the source column yourself4 min

Do not open Q2 yet. Copy this shell, then fill the middle column from the eight rows above, theme by theme. It takes about four minutes and it is the entire skill this session teaches.

ThemeSources you can actually findVerbatim you can actually paste
T1 · do not reopen transcripts??
T2 · the decision is the unit??
T3 · parallel manual note??
T4 · delivered where they work??
T5 · accuracy is the barrier??
T6 · reconstruct from memory??

Two rows will resist you, and they will resist in different ways. In one of them you will find a real quote and then fail to find a second one. In the other you will find yourself reaching for a quote that is nearly about it, and then reaching for another one that is also nearly about it. That second sensation - "this is nearly about it" - is the tell, and it is worth learning to notice, because it is what verification actually feels like from the inside.

LiveQ2 · The reveal - which two, and why4 min

Commit to your two before reading on.

ThemeSourcesVerdict
T1 · do not reopen transcriptsA1, A5, A7, plus A6 for the team behaviourTheme, verified. Also consistent with the whole set: 6 of the 8 never reopen once the meeting ends. Verbatim: "once the call is over I never open the transcript again"
T2 · the decision is the unitA2, A4, A7Theme, verified. Three independent admins, three different phrasings of the same thing. Verbatim: "the thing I need is what we agreed"
T3 · parallel manual noteA1, A3Theme, verified at the floor. Exactly two sources, so it goes forward with its count visible and no more weight than that. The 68% survey figure supports it from a different direction
T4 · delivered where they workA8 onlyOver-generalised. One admin said it, and the theme is written in the plural as a preference of admins in general. Relabel: anecdote, source A8. It is a good hypothesis and it is not a finding
T5 · accuracy is the barriernoneInvented. No snippet says it. Nothing in the material is about accuracy at all
T6 · reconstruct from memoryA4, A7Theme, verified at the floor. Verbatim: "I usually reconstruct that from memory". Closely related to T2 and worth keeping separate, because it names a moment - before a recurring meeting - that T2 does not

Why T5 was so easy to believe. It is a real issue for transcription products in general, and the material has two seams it can hide in: A3 uses the word "trust", and A5 mentions a dispute about what was said. Neither is about accuracy - A3 does not trust that she will find it later, and A5 reopened the transcript to settle a dispute, which is the transcript working. And the one number Cadence has points the other way: accuracy complaints run at 1.4% of meetings, which is the guardrail rather than the problem. A theme with no source can contradict your own instrumentation and still read as the most insightful line in the list.

Why T4 is the more dangerous of the two. T5 dies the moment anybody asks for a quote. T4 has a real quote, from a real admin, and it will survive the first challenge - which means it can travel. And it is not a small claim: "deliver summaries into the tools they already use" is an integration roadmap, sized in engineering quarters, resting on one sentence from one 30-seat agency. Labelled as an anecdote it is genuinely valuable and cheap to test. Promoted to a theme it is a plan.

Self-studyQ3 · Verify a synthesis you already trust4 min

Take a research summary your team is currently acting on - ideally one that supports your carried backlog item, and ideally one you did not write yourself. Do not re-synthesise it. Just add the source column and try to fill it, theme by theme, from the underlying material.

Three outcomes are common and all three are useful. Sometimes every theme checks out, and you now have counts and quotes you can put in the spec in b5. Sometimes a theme turns out to rest on one enthusiastic customer, in which case relabel rather than delete, and notice how much of the current plan was hanging off it. And sometimes the underlying material cannot be found at all - the transcripts are gone, or nobody knows who ran the interviews - which is itself the finding, and it is worth saying plainly, because a theme whose source cannot be located is not evidence any more regardless of how true it once was.

Homework

Try it yourself - this week ◐ 30-40 min total

Source material

Sources covered

Full source map in materials/official-course-map.md. This page covers:

Research-quality literature - sampling, independence of sources, confirmation bias, coherence driftPart 1 + Part 3 · the two-source floor and the independence check
Continuous-discovery synthesis practice - coding before theming, verbatim-backed themesPart 2 + the worked artifact · four steps, and the one that cannot be delegated
AI-assisted knowledge work - where verification cost exceeds the saving, automation biasPart 3 · why the invented theme is always the best-written line
Research ethics and data handling - consent, de-identification at intakePart 2 · the operational rule; consent frameworks are not this course
Interview technique - how to run the eight interviews in the first placeTouched in b2 · this session starts from transcripts that already exist
Delivery: charters, milestones, critical path, status reportingOut of scope by design - learn-ai-project-management
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · You feed in 8 interviews and get back 8 paragraphs, one per interview, all accurate. What have you got?

If the output has as many items as the input, no reduction happened. Eight interviews should produce three to five themes, one of which is probably an anecdote, each carrying a count and a verbatim. Accuracy per paragraph is not the goal and never was.

2 · A theme has one real quote from one real customer behind it. What do you do with it?

Single-source signals are often the earliest sight of something real, so deleting wastes them. But T4 in the worked artifact shows the risk: one sentence from one 30-seat agency, rewritten in the plural, becomes an integration roadmap. Anecdotes get promoted by more research, never by more confidence.

3 · Which failure is harder to catch, and why?

Reading the themes more carefully catches neither, which is the uncomfortable part - you have to go back to the material. And the one with a real quote is the durable one: it passes the verbatim test and fails only the count test, which is exactly the test that gets dropped when a theme is restated in a document three steps from the transcripts.

PM session 3 cheat sheet · pin this

Synthesis isFinding the pattern across accounts. Fewer outputs than inputs, or you summarised.
The flowDe-identify → code → theme → verify. Only the last one cannot be delegated.
Code before themeA code is a claim about one snippet and is cheap to check. A theme is a claim about the set.
Shape A · inventedZero sources. Reads as the most insightful line. Dies when you ask for the verbatim.
Shape B · over-generalisedOne source in the plural. Has a real quote, so it travels. Dies only on the count.
The four rulesCount in the line · one verbatim · one source = anecdote · the PM reads the quotes.
Two review questions"How many people said this?" and "can you show me one saying it?"
The honest costVerifying is always slower, with no speed-up anywhere. Next: b4, themes into a tree.