Before the build, or not at all
b6 made the call: post-meeting summaries for team admins on paid workspaces of 5 to 50 seats, at the cost of the live in-meeting notes sales wanted. The spec from b5 is signable. Engineering could start on Monday. This session is the reason they should not start until Wednesday, and the two days are the cheapest insurance in the whole course.
Everything here is reversible for about a week and irreversible after that. You can choose a different primary metric while the feature is a document. Once it has shipped without the events that produce the metric, your options are to guess, to wait a quarter for the next instrumentation cycle, or to reinterpret whatever numbers you happen to have, which is the readout failure b8 spends a whole session on. Two of those three options are bad and the third is worse.
One primary, and the set around it 6 min live
The leader track argues the rule in a5: one primary per decision, because you can only be judged on one thing at a time. Here we spend it. Below is the actual metric set for Cadence's summaries build, with the rejects named, because the rejects are where the learning is. Measure everything you like. Exactly one of these numbers is the one the team agreed to be wrong about.
LiveLeading or lagging, in the only version that changes what you do4 min▶
Forget the taxonomy. There is one question, and it is about time: if this number moves, when do I find out, and can I still act on it?
- Revisit rate moves in a week. Ship on Monday, read it the following Monday, and if it has not budged you can change the trigger, the surface, the copy, or stop entirely. That is a metric that steers.
- Seat retention moves in a quarter. By the time it says anything, the quarter is over, four other things have shipped, and you cannot separate your effect from theirs. That is a metric that judges.
- Only one of the two can be primary for this build, because the primary is the number you act on. Retention is the thing the company actually cares about, and it is still the wrong choice here, which is the part that feels wrong and is not.
The honest way to hold both is to say so explicitly in the spec: revisit rate is the primary because it is the only one that can steer, retention is the outcome we believe it feeds, and the link between them is a belief rather than a measurement. Write that sentence down. It is what makes the b8 readout possible to have without anybody looking foolish, because you predicted the structure of the argument before you had the data.
LiveThe guardrail, and the property that stops it blaming you3 min▶
The guardrail is the promise you make to the rest of the company: in pursuing our number, we will not damage yours. On Cadence it is transcript accuracy complaints, currently 1.4% of meetings, and the promise is that it does not rise. Support owns the consequences of that number and support did not choose to run this experiment, which is exactly why the promise has value.
The practical part, and the part that only shows up when you write the events: a guardrail needs to be instrumented at least as carefully as the primary, and usually more carefully, because it will be used against you. If your complaint event does not record what the complaint was about, then every pre-existing transcription grumble in the product lands in your guardrail the week you ship. You will spend the readout arguing about attribution instead of about the result, and you will lose, because the number is real and you cannot break it down.
So the complaint event carries an object property with two values, transcript or summary. One property, decided in an afternoon, and it is the difference between "complaints rose to 1.6%" and "transcript complaints held at 1.3% and summary complaints added 0.3%". Those are completely different conversations and the second one is true.
Self-studyWhy three co-equal metrics is the same as none2 min read▶
It is worth being concrete about the failure, because "one primary metric" sounds like a stylistic preference until you watch it break.
Suppose Cadence had named revisit rate, activation and NPS as three co-equal success metrics. Four weeks later revisit rate is up 8 points, activation is flat, and NPS has drifted down by a point. Every person in the room now has a defensible reading. The team that built it leads with revisit rate. The sceptic leads with NPS. The readout becomes a negotiation about which chart goes first, the decision rule cannot fire because it does not know which number it applies to, and the eventual conclusion is that the result was mixed and we should keep going, which is the conclusion that every set of numbers produces when nobody committed in advance.
Instrument all three. Report all three. But name one of them, before the build, as the one the decision hangs on, and let the other two be context. The cost of doing that is that you might be wrong in public, and that cost is the entire product.
The vanity test, and the baseline nobody pulled 6 min live
One question separates a metric from a number that will make everyone feel good for a quarter: if this goes up and nothing else does, do we care? Ask it out loud about every candidate before you choose. It takes ten seconds per metric and it removes most of them, including the ones a drafting tool will hand you first, because the numbers that are easiest to write about are usually the ones that only move in one direction.
LiveThe baseline is three numbers, not one4 min▶
Cadence's revisit rate is 11% and nobody has ever set a target for it. That is not a story about a badly run product; it is the normal state of most product metrics in most companies, and the leader track spends a whole section on why nobody funds the afternoon it takes to fix. Your job here is the afternoon itself, and the thing to know is that pulling "the baseline" means pulling three things:
- The level. 11.0% of recorded meetings, on paid workspaces of 5 to 50 seats. Note the segment. A number pulled across all workspaces would include solo users whose calls are short enough to skim, and it would be a different number measuring a different thing.
- The variance. Twelve weekly readings gave a mean of 11.0% and a range of 9.4% to 12.8%. This is the part everybody skips, and it is what tells you that a post-launch reading of 12.5% is nothing at all. Without it, every small movement is arguable in both directions, forever.
- The shape. Does it dip in the last week of a quarter, rise in January, move when a big account onboards? Twelve readings will not tell you everything, and they will tell you whether you are about to attribute a seasonal pattern to your feature.
Then, and only then, set the target. 30% by week 4 is a real target because there is a real 11% under it and a known band of noise around that. The same 30% with no baseline is a wish that cannot be missed, and a number that cannot be missed will never make anybody change their mind.
LiveThree vanity numbers you will be offered this quarter3 min▶
All three are legitimate operational numbers and all three fail as a primary metric, for the same structural reason: they measure that the feature exists rather than that the problem got smaller.
- Volume of the thing you shipped. Summaries generated per week. It goes from zero to thousands on launch day and the chart is beautiful. It would look identical if every summary were wrong and nobody ever opened one.
- First-touch usage of the thing you shipped. Summary open rate. New things get opened once out of curiosity, so the number is high in week one, decays, and the decay gets read as a problem with the rollout rather than as the shape of novelty. It also cannot really fall below its floor, which means it cannot deliver bad news.
- A composite satisfaction score. NPS or a CSAT average. Moves in a quarter, moves for twelve reasons, and is the single easiest number in the company for a well-written narrative to explain in whichever direction is convenient.
The tell they share: none of them can go down as a result of your feature being bad. A metric that cannot deliver bad news is not an instrument, it is a decoration, and choosing one is how a team spends a quarter and ends it without learning anything about whether the problem moved.
Self-studyWhat to do when the number you want does not exist3 min read▶
The metric Cadence actually wants is time to find a decision: how long it takes a team admin, two days after a call, to answer "what did we agree". That is the real job, it is what the 214 tickets are about, and there is no way to measure it from product events. You cannot instrument a question somebody asks themselves at their desk.
There are three honest responses and one dishonest one. The dishonest one is to pick a proxy and then talk about it as though it were the real thing, which is how a team ends up defending revisit rate as if it were the goal rather than a stand-in for it.
- Pick the closest observable proxy and label it as a proxy, in writing. Revisit rate is the proxy here, and the spec says so in the same sentence that names it. Then the confound is a known limitation rather than a surprise.
- Add a cheap direct instrument alongside it. The in-product pulse question that produces the 68% note-taking number is exactly this: self-reported, weak, and the only direct read available. Weak evidence that you have labelled weak beats strong evidence about the wrong thing.
- Accept that some outcomes are measured by research rather than by analytics, and put the research on the calendar next to the readout. Five interviews at week 6 will tell you things no event ever will, and they cost a fortnight of somebody's attention rather than an engineering cycle.
From the decision down to a property name 5 min live
Here is the whole discipline in one chain. A decision needs a metric; a metric is really a question about behaviour; a question about behaviour needs an event; and an event is only useful if it carries the properties that let you slice it. Engineering can build from the bottom of that chain and cannot build from the top, and nobody tells you which rung you stopped at until the readout, when it is a quarter too late.
LiveProperties are where the readout is won or lost4 min▶
The event names are the easy part and almost everybody gets them roughly right. The properties are where the difference lives, because a property is the only way to ask a follow-up question, and the readout is entirely made of follow-up questions.
Take one property from the spec below: entry_point on transcript_opened, with values search, meeting list, summary link, shared link, notification. It costs an engineer about ten minutes. Now look at what it buys at the readout:
- Did the summary cause the revisit? If revisit rate rose and the growth is all in the summary-link entry point, that is a genuinely different story from the same rise spread evenly across search and the meeting list.
- Is the feature discoverable? If almost nobody arrives via the summary link, the summary might be fine and the link might be invisible, which is a copy and placement fix rather than a strategy change.
- Should phase two exist? If the notification entry point does nothing, the digest email you were planning is a much weaker bet, and you learned that for free.
Without the property, all three of those questions get answered with an opinion. The general rule: for every number you will report, write down the first three questions somebody will ask when they see it, then check that a property exists for each one. That takes five minutes and it is the highest-return five minutes in this session.
LiveWhat AI does well here, and the one thing it must not do4 min▶
This session has an unusually clean split, and it is worth naming precisely because the useful half is genuinely useful.
- Drafting the event spec from the metric definition. Give it the metric, the segment and the denominator, and ask for the candidate events and properties that would produce it. This is real work, it takes fifteen minutes off the task, and every line of the output is checkable against the metric you wrote. Hand it over without hesitation.
- The coverage check, which is the best use on this page. Paste your event list and ask: what questions about this feature could a stakeholder ask that these events cannot answer? It is very good at this, because it is a mechanical comparison rather than a judgement, and it routinely finds the missing property while the spec is still free to change.
- Naming consistency. Whether you are using
summary_viewedorview_summary, whether ids areworkspace_ideverywhere, whether one event uses minutes and another uses seconds. Tedious, mechanical, and worth doing before engineering rather than after. - Never: deciding which number the team is judged on. Ask what the primary metric should be and you will get a fluent answer with a plausible target attached, and it has not seen your baseline, your segment, your company's tolerance for a flat quarter, or the argument you are about to have with sales. The choice of primary metric is a commitment about what you will be wrong about in public. It is the definition of the thing that cannot be delegated.
The second bullet is the one worth building a habit around. It found the gap in the spec below, the team decided to accept it anyway, and b8 shows exactly what accepting it cost.
The event spec for Cadence summaries ★ the whole thing, before the build
Five events, written before engineering started. Read the last column first: every event exists because a specific question needed answering, and any event that cannot name its question should not be in the spec. Then read the gap at the end, because the interesting part of this artifact is the thing it deliberately does not do.
Start from the metric, not from the feature. Write the metric definition first, with its numerator, denominator and segment. Every event below exists to produce one of those three parts, and an event that produces none of them is telemetry rather than instrumentation.
Name the firing condition precisely enough to be wrong. "When the user views the summary" is not a firing condition. "Client-side, on first render of the summary panel for a given meeting, once per workspace per meeting" is, and an engineer can disagree with it, which is the point.
For each event, write the question it answers. One sentence. If you cannot write it, delete the event. This single rule keeps event specs from turning into the 40-event tracking plans that nobody maintains and everybody distrusts.
Run the coverage check before you send it. List the first three questions each stakeholder will ask at the readout, then find the property that answers each one. What you cannot answer becomes a named gap, written down, rather than a surprise.
Then send it to engineering with the metric definitions attached, not separately and not later. The definitions are what let an engineer catch the mistake you made, and they will, roughly a third of the time.
| Event | Fires when | Properties | The question it answers |
|---|---|---|---|
| summary_generated server-side |
a post-meeting summary finishes generating, once per meeting, whether or not anybody ever opens it | meeting_id · workspace_id · seat_band (under_5 / 5_to_50 / over_50) · plan · meeting_minutes · generation_seconds · outcome (ok / no_decisions_found / failed) · cost_cents |
Did the thing we shipped actually run, on which meetings, and at what cost? This is the denominator for everything below, and cost_cents is what tests the 0.02 per meeting estimate against reality. outcome matters more than it looks: a meeting with no decisions in it is not a failure, and if we do not separate the two we will read our own success rate wrong. |
| summary_viewed client-side |
first render of the summary panel for a given meeting, once per workspace per meeting | meeting_id · workspace_id · seat_band · hours_since_meeting_end |
Is anyone looking at it at all, and how soon after the call? Note what is deliberately absent: no user_id and no surface. This event answers "did this workspace open the summary", and it cannot answer "which person" or "from where". That is a known gap and it is recorded below rather than discovered later. |
| transcript_opened client-side · primary metric |
the full transcript view finishes loading for a meeting, every time, not deduplicated | meeting_id · workspace_id · seat_band · user_role (admin / member) · days_since_meeting · entry_point (search / meeting_list / summary_link / shared_link / notification) |
In week 1, did somebody go back to a meeting we summarised, and how did they get there? This is the numerator of the primary metric. days_since_meeting defines week 1 as 0 to 6 inclusive; entry_point is what tells us whether the summary is the doorway or an unrelated bystander. |
| accuracy_complaint_raised mixed · guardrail |
a user flags an inaccuracy in product, or support tags a ticket as an accuracy issue; one event per distinct complaint | meeting_id · workspace_id · seat_band · source (in_product_flag / support_ticket) · object (transcript / summary) · kind (wrong_fact / missing_decision / garbled) |
Are we keeping the promise to support? The guardrail is complaints as a share of meetings, at or below 1.4%. object is the property that stops pre-existing transcription complaints being attributed to summaries, and kind is what turns a breach into a fix instead of an argument. |
| note_taking_survey_answered in-product pulse · secondary |
an admin answers the sampled one-question pulse shown after a meeting; sampled, not shown to everyone | workspace_id · seat_band · user_role · answer (yes / no / sometimes) · wave (pre_launch / week_2 / week_6) |
Are admins still taking their own notes during meetings? 68% today, target 58% by week 6. This is self-report and therefore weak, and the wave property is what makes it comparable at all: the pre-launch wave is the baseline, and without it the week-6 number is a number rather than a change. |
METRIC DEFINITIONS · Cadence post-meeting summaries · agreed before the build
Owner: PM Read date: fixed in the launch note Segment: paid, seat_band = 5_to_50
PRIMARY week-1 transcript revisit rate
numerator distinct meeting_id with >= 1 transcript_opened where days_since_meeting <= 6
denominator distinct meeting_id with a summary_generated outcome of ok, same ISO week
baseline 11.0% · 12 weekly readings, range 9.4% to 12.8%
target 30.0% by week 4
belief going back to the record is the behaviour worth buying. Stated so that
being wrong about it is a finding rather than a reinterpretation.
SECONDARY self-reported in-meeting note-taking
measure share of note_taking_survey_answered with answer = yes, by wave
baseline 68% (pre_launch wave) target 58% by week 6
caveat self-report. Reported as self-report, every time, including in the good case.
GUARDRAIL transcript and summary accuracy complaints
measure distinct accuracy_complaint_raised / distinct meetings recorded, weekly
threshold at or below 1.4% of meetings, split by object
consequence a breach pauses the rollout. It does not trigger a discussion.
COST run cost per meeting
measure sum(cost_cents) / count(summary_generated), weekly
estimate roughly 0.02 per meeting. Tracked, not a guardrail, unless it is 5x out.
KNOWN GAP We cannot answer: "did the person who read the summary decide not to open the
transcript?" summary_viewed is workspace-grained with no user_id and no surface,
and there is no event for a deliberate non-open. Accepted knowingly: the primary
metric does not need it, and adding it would delay the build by a week.
Consequence if this matters at the readout: any statement we make about summary
readers will be about WORKSPACES, never about people. If the argument turns on
whether a summary replaced a transcript read, this line is why we cannot settle it.
The most instructive line in this artifact is the last one, and it is a line about something the team chose not to build. The coverage check found it in about a minute: with these five events you cannot tell whether a person who read the summary then decided they did not need the transcript. Answering that would have meant adding a user_id and a surface to summary_viewed, which meant a client-side change, which meant a week. The team was three weeks into a three-week build. They wrote the gap down and shipped.
That was a defensible call and it was still a call, with a price attached, and the price arrives in b8. When the readout comes in mixed, the single strongest argument in the team's favour is exactly the question this event cannot answer, and because the instrumentation is workspace-grained, the sentence they can write is "71% of workspaces that opened a summary did not open the transcript in the same week" rather than anything about people. It is suggestive, it is not conclusive, and it is not enough to change a conclusion. The next quarter's first instrumentation task is the event they skipped.
The lesson is not that they should have built it. Three weeks is three weeks and the gap was survivable. The lesson is that the difference between a team that writes the gap down and a team that discovers it at the readout is about ninety seconds of coverage checking, and the second team will spend that readout arguing about whether the metric was ever the right one, which is an argument nobody wins.
Try it yourself - this week ◐ 30-40 min total
- Run the vanity test out loud on every metric currently on your team's dashboard: if this goes up and nothing else does, do we care? Count how many fail. The count is the finding, and it is usually more than half.
- Pull a real baseline for your carried backlog item: the level, the variance over at least eight readings, and the segment it is scoped to. Then check that your target sits outside the historical range, because if it does not, you have described a normal week.
- Write the guardrail as a sentence addressed to another team, with a threshold rather than a direction, and check that it has the one property that stops it blaming you for something that predates your feature.
- Write the event spec for your item before engineering starts: names, firing conditions, properties, and the question each event answers. Delete every event whose question you cannot write in one sentence.
- Run the coverage check: list the first three questions each stakeholder will ask at the readout, find the property that answers each, and write down the gaps you are choosing to accept. Bring the decision rule for that metric to b8.
Sources covered
Full source map in materials/official-course-map.md. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · Seat retention is what the business actually cares about. Why is week-1 revisit rate the primary metric for this build instead?
The primary is the number you act on, and you can only act on a number that arrives while there is still something to change. Retention stays in the spec as the lagging outcome you believe revisit rate feeds, and the link between them is written down as a belief so that being wrong about it is a finding rather than an argument.
2 · "Summaries generated per week" is proposed as the success metric. What is wrong with it?
A metric that cannot go down as a result of your feature being bad is a decoration rather than an instrument. Volume of the thing you shipped, and first-touch usage of the thing you shipped, both measure that the feature exists rather than that the problem got smaller.
3 · Which of these is the one job in this session that must never be handed over?
Drafting the events from a metric you wrote, and mechanically checking the list for coverage gaps, are both production work with a cheap verification step - and the coverage check is the single best use on the page. Choosing the primary metric is a commitment about what you will be wrong about in public, and it has never been anything else.