Why this session exists at all
a1 drew the line between the half of the job that got cheap and the half that did not. This session takes one consequence of that line and follows it all the way down, because it is the consequence that will cost you a quarter if you miss it. Research summaries are production work. Production work is now free. And research summaries are the input to almost every product decision you approve.
You are not going to learn interviewing technique here, and you are not going to audit anybody's spreadsheet. Your PMs do that work; the practitioner track spends a whole session making an invented theme appear on purpose so they learn to catch it. Your job is narrower and more powerful: to be the person in the room who reliably asks where a claim came from, in the same words, every time, until nobody brings you a claim without its provenance attached. That is a change to one sentence of your behaviour, and it changes what your organisation produces.
Where a leader would expect a delivery question - can we resource the research, when does it land - that is a planning matter and belongs to learn-ai-project-management. This course stays on whether you should believe the finding.
The bar moved, and nobody announced it 5 min live
Think about what a research readout used to tell you before you read a word of it. Somebody had recruited participants, sat through the calls, read the transcripts, argued with a colleague about what the themes were, and then written it up. The document was not the evidence. The document was the receipt for the evidence, and it was a fairly reliable receipt because forging it took nearly as long as doing the work.
LiveThe three sentences that should now stop you4 min▶
These used to be shorthand for work somebody had done. Now they are shorthand for nothing, and they arrive in beautifully formatted documents:
- "Users are telling us..." Which users, how many, telling whom, when? The construction hides the count, and it hides it whether the count is forty or one. When the count is good, people say it, because a good count is the strongest thing in the sentence.
- "The research shows..." Shows is doing a lot of work here. Research does not show things; specific people said or did specific things, and somebody interpreted them. The passive form removes the interpreter, which is exactly the person you need to talk to.
- "There is a clear theme around..." Clear to whom, and derived from how many sources? "Clear" is a claim about the reader's confidence rather than about the data, and a fluent summary is extremely good at producing confidence.
Self-studyWhy your own judgement is not the defence3 min read▶
The tempting response to all this is "I will read more carefully". It does not work, for three reasons worth being honest about:
- You cannot detect absence by reading. A theme with no source behind it does not read differently from a theme with eight. There is nothing on the page to notice, which is why the fix has to be a required field rather than a sharper eye.
- You audit selectively, and you audit in the wrong direction. Everyone interrogates findings that contradict what they expected and waves through findings that confirm it. The findings that confirm your view are the expensive ones to get wrong, because they are the ones you act on immediately.
- Careful reading does not scale and it changes how you are treated. A leader who sometimes interrogates deeply and sometimes does not teaches their team to guess which mood they are in. A leader who asks the same two questions every single time teaches their team to answer them in advance, which is free.
So the standard has to be a requirement, uniformly applied, boring, and stated the same way every time. Boring is the feature.
Three grades of evidence 7 min live
Every claim that reaches your desk is one of three things, and telling them apart takes about four seconds once you have the categories. The reason this matters is not tidiness. It is that a document flattens all three into the same typeface, the same bullet shape and the same confident tone, so a belief and a measurement look identical on the page.
| Grade | What it is | What it is worth | The one question that tests it |
|---|---|---|---|
| 1 · Counted behaviour | What people did, measured, re-pullable by somebody else | You can act on it, once you know what was counted and over what window | "Where does this number come from, and can somebody re-pull it?" |
| 2 · Reported behaviour | What people said they do, want, or would pay for | A hypothesis. Weight depends entirely on how many people and who chose them | "How many people, and were they asked or did they raise it themselves?" |
| 3 · Asserted belief | What somebody thinks is true, including when they are senior and confident | An agenda item. Legitimate input, never a reason on its own | "What would we see in the data if this were true?" |
LiveGrading the three Cadence asks, live4 min▶
Three people want three things from the same product. Grade each ask before you discuss any of them, and notice how much of the argument evaporates:
- Support wants post-meeting summaries. 214 tickets in three months tagged "cannot find what was decided", the second most common tag. That is grade 1, and it is re-pullable, which is what makes it grade 1 rather than a good story. The follow-up question is what the tag actually means and whether the 214 came from a handful of accounts. Not a reason to discount it, a reason to ask.
- Sales wants live in-meeting notes. "Every demo asks for it" is grade 3: an assertion, from someone with an incentive, using an absolute quantifier that is almost certainly not literal. The three lost deals sound like grade 1 because they are a number, but they are three narratives about why a deal was lost, told by the person who lost it. Treat as grade 2 at best, n equals three, non-random.
- The CEO wants an agent. Grade 3, offered with no evidence, and the hardest of the three to say no to. Grading it does not make it go away and should not: a CEO's belief about where the market is going is a real input. It is just not evidence about your users, and the honest move is to say which one you are treating it as.
Self-studyGrade 1 is not automatically the strongest3 min read▶
The grades rank how much a claim can be checked, not how much it matters. Three ways grade 1 misleads a leader who has just learned to love numbers:
- A number with no target is not a finding. Cadence's week-1 transcript revisit rate is 11% of recorded meetings, and no target was ever set for it. Is 11% bad? Nobody in the company can answer that, which means the number currently supports any argument that wants it. a5 is where targets and baselines get their own session.
- A count is a count of what was tagged, not of what happened. 214 tickets tagged "cannot find what was decided" is a fact about a taxonomy that support staff apply under time pressure. It is still the best evidence in the room. It is not a fact about 214 distinct problems.
- The strongest grade-2 evidence beats weak grade 1. Six of eight interviewed admins saying they never reopen a transcript is grade 2 with a tiny sample, and it is more useful than most dashboards, because it tells you why the 11% is 11%. Grades tell you what question to ask next, not which claim wins.
So the honest formulation is: grade 1 can be checked, grade 2 explains, grade 3 sets the agenda. A decision usually needs the first two, and gets presented with the third.
Provenance, and the two failures 7 min live
Grading tells you how much weight a claim can carry. Provenance is what you actually require in order to grade it, and it is three things: how many sources, as a number; at least one verbatim quote you can read yourself; and who collected it, and when. If those three are present, you can grade any claim in seconds. If they are absent, no amount of careful reading will help you, because there is nothing on the page to read.
LiveFailure one: the theme with nothing behind it4 min▶
Six themes in a synthesis. Five are drawn from things people actually said. The sixth is a highly plausible statement about note-taking behaviour that no participant made, produced because it fits the shape of the other five and reads well beside them.
- It is usually the best-written theme in the document. Themes grounded in messy real quotes are messy. An ungrounded one has no awkward detail to accommodate, so it comes out clean, general and quotable, and it is often the one that gets repeated in the next meeting.
- It survives review because it is not wrong. This is the part leaders find hardest to internalise. It is not a false statement. It is plausible, sensible, probably even true in general. It is simply not your evidence about your users, and a decision built on it is a decision built on a general truth about humans.
- The only defence is the source count. Nothing about the sentence gives it away. Ask which participants said it, and there is either an answer or there is not. This is why the count has to be a required field rather than something you ask for when a line feels off.
Your PMs meet this one deliberately in the practitioner track, with real interview snippets and a synthesis that contains a planted invented theme. You do not need to run that exercise. You need to be the reason it never reaches a decision.
LiveFailure two: one quote wearing a plural4 min▶
The second failure is more common than the first and much harder to see, because there genuinely is a source. One participant said something vivid and specific. The synthesis promotes it into a general claim about the segment.
- What it looks like. One admin says "I stopped opening transcripts after the first week, I just ask the person". The theme reads: "Admins abandon transcripts within a week and prefer to ask colleagues directly." Every word is defensible. The plural is doing enormous unearned work, and the sample size is one.
- Why summarisation produces it. Turning specifics into generalities is exactly what good summarising is. It is the desired behaviour, applied to a place where it is dangerous, which is why the person doing the summarising cannot be relied on to flag it.
- The tell, when there is one. A vivid theme with exactly one supporting quote. When six people said something, a synthesis usually offers two or three quotes because they are available and they reinforce each other. One quote for a general claim is worth a question, every time.
Self-studyWhat to do when provenance is genuinely thin3 min read▶
Sometimes the honest answer is "two people, and here is one of them". That is not a failure and you should not punish it, or you will teach your team to hide n.
- Thin evidence, stated as thin, is usable. "Two admins said this, both on the team plan, and it matches the revisit rate" is a perfectly good basis for a small cheap experiment. It is a terrible basis for a quarter of engineering, and the difference is size of commitment rather than quality of evidence.
- Match the size of the bet to the grade of the evidence. This is the whole practical use of the grading. Grade 3 buys an agenda slot. Grade 2 with small n buys a week of research or a cheap test. Grade 1 with a baseline buys a quarter. a4 and a5 make this concrete.
- Reward the person who says "we do not know". If declaring uncertainty is punished, uncertainty stops being declared, and you lose the only signal that tells you when to slow down. This is the same cultural question a5 raises about readouts that admit the bet failed, and it has the same answer: you get what you visibly reward.
The provenance drill ★ run this in your next product review
One document, two questions, timed. This works in the room today and it works unchanged in your next product review, which is the point: it is not an exercise, it is the behaviour, rehearsed once so that it does not feel confrontational the first time you do it for real.
Everyone puts one real research document on the table. A themes list, an insight summary, a customer readout. Ideally the one you brought because you believe it - the confidence is what makes the exercise land.
Pick the single theme in it that most influenced a decision. Not the weakest one, the most load-bearing one. Underline it.
Ask the two questions out loud, in order, and time the answers. "How many people said that?" then "Can I read one of them, in their own words?" Say nothing else. Do not fill the silence, and do not soften it into three questions.
Write the grade next to the theme. Counted, reported, or asserted. Then write the source count. If you cannot write a number, write a question mark, and notice how that feels next to a decision you already approved.
Say what you will now require by default, in one sentence, out loud. Commit to the wording in the room so that the first time you use it for real, it sounds like a standard rather than an accusation.
LiveThe prompt, word for worduse this exact wording▶
What a good answer sounds like. "Six of the eight admins we interviewed. Here is one: 'I stopped opening them after the first week.' Priya ran the interviews in March, all paid workspaces between 5 and 50 seats." A count, a quote, a collector, a date, and a scope. Delivered in under fifteen seconds, because the person answering has it to hand. Note that a good answer volunteers the sample frame without being asked - who was not in the eight is usually the more interesting half.
The failure mode to listen for. The unquantified plural, then the delay. "Several admins", "users are saying", "it came up a lot" - followed by "let me check and come back to you". That is not dishonesty and you should not treat it as dishonesty; it is what happens when the count was never a required field. The second failure mode is quieter: a single quote produced smoothly for a theme phrased in the plural. Listen for the mismatch between the quote and the claim, and ask who else said it. If the answer is a pause, n is one.
Say the framing out loud first. "I am going to ask this about everything from now on, including things I agree with, so it is not about this document." Skip that sentence and the room learns that provenance questions signal distrust, which is exactly the wrong lesson and it will cost you the next three findings.
LiveSecond round: run it on something you agree with4 min▶
What a good answer sounds like. The same shape as before, plus a bit of discomfort. The useful outcome is somebody saying "actually I never checked that one, because it matched what I expected". That sentence is the whole exercise, and it is worth more than any framework on this page.
The failure mode to listen for. The reflex defence: "that one is obviously right". Obviousness is a fact about the reader, not about the evidence, and it is precisely the condition under which nobody checks. Watch also for the room grading confirming findings more generously than contradicting ones. Say it out loud when you see it, including when you are the one doing it, which you will be.
Cadence, four claims, four different grades. On the table: "214 support tickets in three months tagged cannot find what was decided, the second most common tag" - grade 1, re-pullable, and the strongest thing in the room. "Week-1 transcript revisit rate is 11%" - grade 1, and useless on its own until somebody sets a target, because nobody in the company can say whether 11% is bad. "68% of surveyed admins still take their own notes during meetings" - grade 2, and the most interesting line available, because it says the product's central promise is not being trusted by the people paying for it. "Every demo asks for it" - grade 3, from someone with a quota, using an absolute quantifier. All four appeared in the same document, in the same font, as four bullets of apparently equal weight. Nothing in the formatting distinguished the re-pullable count from the sales assertion, and the person who wrote it was not being careless: flattening is what a well-formatted document does.
Questions to ask your PMs the whole point of the session
Seven questions. The first two are the ones that matter and you should ask them until they become a joke about you, because a leader with one predictable question changes behaviour faster than a leader with a policy.
- "How many people said that?" Works because it is unanswerable with a plural. It converts every soft synthesis claim into either a number or a visible gap, and it takes four words. If you only adopt one thing from this session, adopt this.
- "Can I read one of them, in their own words?" Works because it proves the source exists and shows you the distance between what one person said and what the theme claims. That distance is where over-generalisation lives, and no other question exposes it as cheaply.
- "Who collected this, and when?" Works because it makes the evidence attributable to a person rather than to a document. It also surfaces the stale finding: research from fourteen months ago is being used as though it were current more often than anyone admits.
- "Is this what they did, or what they said, or what somebody thinks?" Works because it does the grading out loud without jargon, and it invites an honest answer instead of a defence. Most people know which grade their claim is, and will tell you if asked plainly.
- "Who was not in the sample?" Works because sample framing is where friendly-customer bias lives, and nobody volunteers it. Eight interviews with your most engaged admins pass every count-based check and can still point you in the wrong direction.
- "What would we see in the data if this were true?" Works on grade 3 specifically, including when grade 3 arrives from the CEO. It converts a belief into a testable claim without contradicting anybody, which is the only way that conversation goes well.
- "Which of these findings did you check, and which did you accept?" Works because it acknowledges that nobody checks everything, so it gets an honest answer rather than a defensive one. The findings that were accepted rather than checked are almost always the ones that confirmed the plan.
Five things to do before a3 ◐ conversations, not documents
- Ask "how many people said that, and can I read one of them" in every review you attend this week, without exception, including when you agree with the finding. Count how many times you get a number in under fifteen seconds.
- Tell your team, once, in a room, what you will require from now on: a source count, one verbatim, and a collector with a date. Frame it as a standard rather than a response to something. If you frame it as a response, somebody will assume they are in trouble.
- Take the last significant product decision you approved and grade the evidence behind it: counted, reported, or asserted. Do this alone. If it was grade 3, that is worth knowing before a4 asks you to rank things.
- Find the oldest research finding your team still repeats as current, and ask when it was collected. There is one, it is older than you think, and the conversation is usually short and useful.
- Praise somebody publicly for saying "we do not know that yet". Once is enough to be noticed, and it is the cheapest cultural purchase available to you this quarter.
Sources covered
Full source map in materials/official-course-map.md. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · A well-structured six-page research synthesis lands on your desk. What does its existence now tell you about whether research was done?
The correlation between polish and effort held for your entire career because writing three fluent pages of synthesis used to cost an afternoon. It now costs under a minute. Nothing about the document distinguishes the two cases, which is why the fix is a required source count rather than a sharper eye.
2 · Sales says "every demo asks for live notes" and offers three lost deals. Support offers 214 tickets in three months tagged "cannot find what was decided". How should you treat them?
Grading is not a way to win an argument with sales - it is what makes two incomparable claims comparable, which is the only condition under which choosing is possible. The three deals look like grade 1 because they are a number, but they are narratives with n equal to three and an interested narrator.
3 · A synthesis contains a vivid, well-written theme about admin behaviour with exactly one supporting quote. What is the most useful next question?
When six people said something, a synthesis usually offers two or three quotes because they are available and reinforce each other. One quote carrying a plural claim is the signature of the second failure mode, and the verbatim requirement earns its keep here by showing you the distance between what one person said and what the theme asserts.