learn-ai-project-management-with-phoebe / Leader session 6 of 6
Learn AI + Project Management with Phoebe · Leader track · Session 6 of 6

Rolling it out to a PMO: one programme, one artifact, six weeks, one volunteer

Everything in this track works on one desk. Tonight you make it work on twenty, which is a different problem with different failure modes - most of them human. You will design a six-week pilot, set an acceptance bar you can sample rather than police, pick metrics your PMs cannot game in an afternoon, and handle the part nobody puts on a slide: adopting a new tool with people who have been burned by the last three.

🟠 Leader track Heads of delivery PMO leads Rollout plan Finale
0-3 · Welcome 3-20 · Pilot, quality bar, metrics 20-42 · Charter + measurement plan 42-45 · Close
Part 0

Where we are, and where this ends

Five sessions in. a1 drew the line between drafting and judgment. a2 set the governed boundary - what goes in a workspace, what never does, who connects what. a3 gave you a review rubric for AI-drafted plans, a4 did the same for risk, dependencies and escalation, and a5 built the reporting cadence and the six checks. All of it has been about one programme, one desk, one PM. This is the finale, and it changes scale: how you take something that works for one person and put it into a function of fifteen or fifty without breaking trust, quality or anyone's Friday. The answer is not a rollout plan. It is a pilot with a stop condition.

Live - presented in session Self-study - read after class ★ Try it now prompt Official sources covered
★ What you walk out with today A six-week pilot charter you could start on Monday, an acceptance bar based on the six checks with a sampling rate instead of a police force, a metric set that is honest rather than flattering, an adoption approach that survives PMs who have been burned before, and one clear answer to "what do I do first".
Part 1 · covers PMI's governance principle + adoption reporting

The pilot: one of everything 5 min live

The instinct is to roll it out to everyone, because the value looks obvious and the licence is already paid for. Resist it. A wide rollout gives you twenty programmes changing at once, no baseline, no control, no way to tell whether anything improved, and a support load that lands entirely on you. Worse, it gives every sceptic in the function a story to tell by week three. One programme, one artifact, six weeks, one PM who volunteered.

ONE PROGRAMME · ONE ARTIFACT · ONE VOLUNTEER · SIX WEEKS Week 1 score today's report, 3 of 6 Week 2 load the pack, agree the rule Week 3 first drafted report sent Week 4 sampled review, fix what fails Week 5 2nd PM joins if it works Week 6 score again, go or no-go Everything outside the pilot stays exactly as it was. That is what makes week 6 readable. If nothing moved by then you stop, and a clean stop is a real result, not a failure. Go means score up, hours down, and the volunteer wants to keep it. Two of three is a not yet.
🔍 Click to zoom - six weeks, one artifact, and a go/no-go that can honestly say no
LiveWhy the status report is almost always the right first artifact3 min

People reach for the exciting artifact - the plan, the business case, the risk model. Pick the boring one. The weekly status report wins on three properties that nothing else in the PM canon has at once.

  • Highest frequency. It happens every week on every programme, so six weeks gives you six real data points per PM instead of one. A charter pilot gives you one event and no trend.
  • Clearest rubric. The six checks are objective enough that two reviewers agree: is the RAG rule written in the report, does every action have one name and one date, is there a traceable id behind each claim. You cannot score a charter that cleanly.
  • Lowest blast radius. A bad drafted report is caught on Wednesday by the person who signs it on Friday. A bad drafted plan gets committed to and lives for a year. Start where the mistakes are cheap and visible.

Public reporting says the same thing from the other direction: roughly half of project professionals say AI already touches their work, about a fifth have real hands-on experience, and status reporting is consistently the first place it lands. You are not choosing the frontier here. You are choosing the place your function is already drifting, and putting a rubric under it before it becomes habit.

Real world

The pilot that picked the wrong artifact. A PMO ran its first AI pilot on business cases, because that was where the executive attention was. Six weeks produced two business cases, both heavily edited, and no way to say whether anything had improved - the sample size was two and the quality bar was "the exec liked it". They restarted on the weekly report, got thirty-six reports in six weeks across three PMs, and could show a score moving from 3 of 6 to 5 of 6 with the exact check that lagged. Same six weeks, entirely different evidence.

Self-studyChoosing the volunteer, and why it must be a volunteer2 min read

The single most common pilot failure is a nominated participant. Nominated people produce compliant behaviour, compliant feedback, and no information. You want someone who put their hand up, because only they will tell you the part that did not work. Pick a credible PM rather than the keenest one - if it succeeds, the sceptics have to believe the result, which the department enthusiast cannot deliver.

Pick a programme that is mid-flight and slightly messy, because a pristine one proves nothing. Northwind at week 9 - a gate slipped from 10 August to 24 August, a vendor gone quiet, quality at 96.4% against a 99% gate - is exactly the sort of week worth testing on. Do not pick a programme in crisis: nobody learns anything in a fire, and the tool gets the blame for it. And protect your volunteer explicitly, out loud, in front of others: if the pilot produces nothing, that is a finding and not a performance issue. Otherwise they will find something, and you will have bought yourself a lie.

Part 2 · the acceptance test and how to sample it

The quality bar, sampled not policed 4 min live

You already have the acceptance test: the six checks from a5. What you need now is a way to apply it that does not turn your PMO into a review queue. Reviewing everything is how quality programmes die - it costs more than the thing it protects, it makes the reviewer the bottleneck, and it signals to every PM that you do not trust them. Sample instead.

The checkWhat passing looks like, concretely
Status by a written RAG ruleThe colour and the rule appear on the same page, in the report itself
The delta since last weekAt least one number stated twice: this week and last. "24 August, was 10 August"
One named owner and one date per actionA person's name, not a team. "Priya N., 6 August". Any TBC older than two weeks fails
Decisions taken, loggedAn id, a date and a decider. "D-14, 29 July, Elena V." Or an explicit "none this week"
Evidence behind every claimA ticket id, a metric with its target, or a named document. NWD-412. 96.4% against 99%
The askThe decision, the options with costs, and the date after which the option lapses
LiveHow to sample, and what to do when one fails3 min

Pick a fixed share of artifacts, review those properly, and leave the rest alone. The exact rate matters far less than the fact that it is fixed, published and unpredictable in its selection.

  • Start around one in five, chosen at random, for the first quarter. Then drop it as scores stabilise. A share that never falls is a police force wearing a sampling badge.
  • Review the artifact, never the person, and publish only the aggregate. The finding is "this report scored 4 of 6, the ask and the delta are missing", not "this PM is weak on reporting". The function sees "we are averaging 4.6 and the ask is our weakest check", and nobody sees whose report was whose.
  • Sample the gaps blocks too. A report that scores 6 of 6 with an empty gaps block every single week is not a triumph, it is a PM who has learned to stop declaring what they could not source.

When a sampled artifact fails, the response is graded and it never starts with the PM:

  • First failure of a check: fix the shared prompt or the house rules. Most failures are systemic - if four PMs all miss the ask, the template is missing a section, not the people. Second time, same PM, same check: fifteen minutes with them, on the check, not on their competence.
  • A failure that reached a sponsor with a wrong number in it: that is a real incident. Reconstruct how the number got in, whether the signing human read it, and whether the pack was stale. Almost always it is a stale pack or a claim with no id, and both are fixable in the cadence.
Real world

The check that was the template's fault. A PMO sampled twelve reports in month one and found the ask missing in nine of them. The obvious read was that PMs were avoiding hard conversations. The actual cause was that the shared report prompt listed the ask last, after a length limit, so it kept getting truncated. One line moved, and the next sample had the ask in eleven of twelve. Nobody had been avoiding anything, and if the lead had opened with a conversation about courage she would have insulted nine people for a formatting bug.

Self-studyWhere the bar sits inside your existing governance2 min read

Do not build a parallel structure. PMI's 2026 standard asks for clear structures, roles, authority and guardrails inside existing governance, with real intervention triggers - not a separate AI committee with its own reporting line and no power. In practice that is four small amendments to documents you already have.

The house rules file gains the six checks and the RAG rule. The PMO assurance routine gains the sampling rate, as one more artifact type rather than a new process. The steerco pack gains one line: the aggregate report score and the weakest check this month, which makes quality visible without naming anyone. And the escalation path gains one trigger - a wrong figure reaching a sponsor is an incident and is logged like any other, which is what turns oversight from a claim into a mechanism. Four amendments beat a new framework nobody reads, and it is the version your risk and audit colleagues will actually accept.

Part 3 · measuring the thing, not the enthusiasm

Metrics that are honest 4 min live

Someone above you will ask for a number by week four. What you give them determines what your function optimises for the next two years, so choose it before you are asked. The test is simple: could a PM move this metric in one afternoon without the programme being any better run? If yes, it is a vanity metric and it will be gamed, not out of cynicism but because people move what you measure.

Measure these Never these · Hours moved off artifact labour, per PM per week · Report score against the six-check rubric · Findings that keep recurring in the gaps block · Decisions per steerco, and slip-to-escalation days · Prompts run, messages sent, tokens burned · AI adoption %, however you define it · Seats licensed against seats logged in · Time saved, self-reported, with no baseline The left lane changes how the programme runs. The right lane changes only the slide about it.
🔍 Click to zoom - two lanes: metrics that move delivery, and metrics that move a slide
LiveThe five honest metrics, and how to actually collect them3 min
  • Hours moved off artifact labour. Ask each pilot PM for one number in week 1: how long the weekly pack takes them, door to door. Ask again in week 6. It is self-reported and imperfect, and it is still the number your finance director will care about most. Be honest that it is an estimate, and never annualise it into a headcount saving unless you intend to make one.
  • Report score against the rubric. Out of six, from your sample, averaged across the function. This is the quality half, and it is the number that stops the hours metric from turning into "fast and worse".
  • Recurring findings in the gaps block. Count how often the same gap appears across three weeks. A gap that keeps returning is a structural hole in the programme - an unowned workstream, a metric nobody produces - and finding those is arguably the biggest prize in this whole track.
  • Decisions reached per steerco, and days from slip detected to escalation raised. Count the decisions: if the ask is working, that number goes up and the meeting gets shorter. The escalation lag matters most and gets measured least - if your Tuesday sweep is working it shrinks, so baseline it from three past slips before you start or you will never be able to claim it.
  • And two rules for whatever you report upward. Report the shape, not false precision: "the weekly pack went from about half a day to about an hour, across three PMs, over six weeks" is defensible, "41% efficiency gain" is not, and the first person who tests it will find that out in front of you. Always pair an hours number with the quality number in the same sentence, so nobody can quote half of it.
Real world

The adoption percentage that measured nothing. A delivery function reported "82% AI adoption" to its board, defined as PMs who had opened the tool that month. The board was delighted. Six months later someone asked what had improved and there was no answer, because nothing had been baselined and the metric had only ever measured logging in. The team quietly rebuilt around report score and days-to-escalation, both of which were worse than anyone had assumed and both of which then improved for real. The second set of numbers was smaller, later, and worth something.

Part 4 · the human half

Adoption with PMs who have been burned 3 min live

Your PMs have survived a portfolio tool, a methodology change, a reorganisation and at least one thing that was going to transform delivery and then quietly disappeared. When you arrive with the next one, the resistance you meet is not irrational. It is memory. PMI's people-and-culture principle sits alongside the technical ones for exactly this reason, and it carries the same weight as governance and risk.

LiveThree things to do, and one thing to say out loud3 min
  • Opt-in beats mandate, for longer than feels comfortable. Mandates produce compliance, and compliance produces artifacts that technically pass and teach you nothing. Keep it voluntary through the pilot and the first wave after it. If the thing works, the second wave asks to join.
  • Show one PM's Friday afternoon coming back. Not a slide about productivity. A named colleague saying what they did instead: went home, or spent it on the vendor who had gone quiet. That story travels through a PMO faster than any number you will ever produce.
  • Never measure individuals against each other, and make the failed pilot safe. The moment a leaderboard of report scores exists, the scores become the work and the gaps blocks go empty, so aggregate up, always. And say the stop condition before you start, then mean it - a pilot that can only succeed is not a pilot, and everyone in the room knows it.

And the thing to say out loud, in the first session, in these words or close to them: nobody's accountability moved. The plan is still theirs. The date they commit is still theirs. The report with their name on it is still theirs. The only thing that changed is how long the first draft takes. That sentence does more for adoption than the demo, because the fear underneath the resistance is usually not about the tool. It is about whether the thing they are accountable for is about to be decided by something they cannot argue with.

Real world

The sceptic who became the proof. The most vocal doubter in one PMO refused the pilot and said so in a team meeting. The lead did not push, and did not exclude him either - he stayed on the review sample, reading other people's drafted reports and scoring them. Eight weeks later he asked to join, and his stated reason was that he had spent two months looking for the failure and had mostly found reports that were harder to argue with than his own. He then became the person other PMs asked, which no mandate would ever have produced. Being allowed to say no is what made his yes worth anything.

Demo 1 of 2

Write the pilot charter ★ 12 min · everyone builds

One page. If it does not fit on one page you are piloting too much. Fill this in for your own function now, with real names in it, and you will leave tonight with something you could start on Monday.

Name the programme - one, mid-flight, slightly messy, not in crisis - and the artifact, which is the weekly status report unless you have a genuinely better reason: highest frequency, clearest rubric, lowest blast radius.

Name the volunteer, and the person covering them. If nobody volunteered, that is your first finding and the pilot has not started yet.

Write the quality bar - the six checks and a sampling rate you will actually sustain - then the go/no-go criteria, before week 1, including the version where you stop. Read them out to the volunteer.

★ Try it now - the six-week pilot charter, one pagePILOT: AI-assisted weekly status reporting PROGRAMME: .................. SPONSOR: .................. VOLUNTEER PM: .................. COVER: .................. RUNS: week 1 to week 6, starting .......... IN SCOPE one artifact: the weekly exec status report, drafted in the workspace, on the a5 cadence - Mon refresh, Tue sweep, Wed draft + gaps, Thu decide, Fri sign OUT OF SCOPE every other programme, artifact and PM. Nothing else changes. QUALITY BAR - The six checks: written RAG rule · delta since last week · one named owner and one date per action · decisions logged · evidence behind every claim · the ask. - Sampling: 1 in 5 reports reviewed at random. Artifact reviewed, never the person. - Aggregate score published to the function. Individual scores published to nobody. GO / NO-GO AT WEEK 6 - all three, or it is a not yet 1) Report score up from the week 1 baseline of ...... of 6. 2) Time on the weekly pack down from ...... hours, self-reported, same PM. 3) The volunteer wants to keep it, unprompted, in their own words. WE STOP IF score has not moved by week 6, or a wrong figure reached a sponsor twice. STOPPING IS A RESULT. It is written here so nobody has to be brave in week 6.
The charter that guarantees an unreadable week 6PILOT: roll out AI across the PMO SCOPE: all 14 PMs, all artifacts, from Monday SUCCESS: improved efficiency and better quality reporting REVIEW: at the end
Why the second one cannot be evaluated No baseline, no control, no artifact boundary, no named human, no stop condition, and a success criterion that cannot be false. At the end you will have opinions and a support queue, and whoever shouts loudest in week 5 will decide the outcome. The narrow charter is not caution - it is the only version that produces evidence.
Demo 2 of 2

The six-week measurement plan ★ 10 min · build your own

A measurement plan written in week 1 is evidence. The same numbers assembled in week 6 are a story you told yourself. The difference is entirely the baseline, and it takes about twenty minutes to capture before you start.

Capture the baseline this week, before anything changes. Score three recent reports out of six. Ask the PM how long the weekly pack takes, door to door. Pull three past slips and count the days from detection to escalation.

Write down what each number is, exactly, and who collects it. "Report score" means the six checks scored by the PMO lead from a random sample. "Hours" means one self-reported number from one PM, and you will say so when you report it. Then set the rhythm: score weekly, hours in weeks 1 and 6 only, decisions counted every steerco, escalation lag baselined once.

Write the result that would make you stop, in a number - not "if it does not work" but "if the score has not moved from 3 of 6 by week 6". Then write the honest sentence you will say upward, in advance: shape rather than false precision, always paired with the quality number.

★ The measurement plan - fill it before week 1METRIC BASELINE (wk 1) COLLECTED BY RHYTHM Report score, out of 6 ...... of 6 PMO lead weekly, 1-in-5 sample Hours on the weekly pack ...... hrs the PM, self-rep wk 1 and wk 6 only Recurring gaps-block findings ...... repeats PMO lead counted at wk 3 and wk 6 Decisions reached per steerco ...... chair every steerco Days, slip detected to escalated ...... days PMO lead baselined from 3 past slips WE STOP IF the score has not moved from the baseline by week 6, OR a wrong figure reaches a sponsor twice, OR the volunteer would rather go back and says so. WHAT I WILL SAY UPWARD AT WEEK 6 - shape, not false precision "Across three PMs over six weeks the weekly pack went from about half a day to about an hour, and the report score went from 3 of 6 to 5 of 6 on a 1-in-5 sample. The hours figure is self-reported. The weakest check is still the ask." WHAT I WILL NOT SAY: a percentage efficiency gain, an annualised saving, an adoption rate.
Real world

The twenty minutes that made the case. A PMO lead spent week 1 doing nothing but baselining: three reports scored, one PM's honest estimate of the pack, three past slips counted back to detection. It felt like wasted time and one person said so. At week 6 she had a before and an after for every claim, including one that got worse - the gaps blocks got longer, because the PMs were finally declaring what they could not source. She reported that too, and it was the number the sponsor found most convincing, because nobody presents their own bad number unless the rest is real.

Closing the leader track

What to do on Monday 2 min live

Six sessions, and the whole track collapses into a short list. None of it needs a budget, a business case or a vendor conversation to start.

Where to go next, and who to send This track deliberately stopped at the leadership half. If you want to build the artifacts yourself - the workspace, the charter, the milestone architecture, the dependency map, the self-maintaining risk register, the scored status report - the PM track starts at b1, your programme workspace, and runs ten sessions on the same Northwind programme. Send your volunteer to b1 and b6 before the pilot starts. If you do one of the two, do b6: it scores the same six checks live, and watching a report climb from 24 to 100 as each check goes in does more for belief than any policy you could write.

The three rails have not changed since a1, and they will outlast every product name in this course. Context beats prompting. The dates come from the team. A named human signs it. Everything else is technique, and technique ages in months.

Homework

Try it yourself - this week ◐ 30-45 min total

Source material

Official sources covered

Taught from PMI's public standards and AI guidance, and the Northwind case carried through both tracks. Certification, the full normative text of the AI standard, and vendor product training stay with their official providers. This session covers:

PMI - Standard for AI in Portfolio, Program and Project Management (2026)Parts 2 + 4 · people and culture, and governance inside existing structures
PMI - AI in project management, adoption reportingParts 1 + 3 · where AI lands first, and the practitioner experience gap
PM track b9-b10 - connectors, skills, and the full delivery weekPart 1 · what a PM builds once the pilot goes past one artifact
PMBOK Guide 7 - measurement performance domainPart 3 · baselines, and why a metric without one is a story
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · Why is the weekly status report almost always the right first artifact for a pilot?

A charter pilot gives you one event and no trend. Thirty-six reports over six weeks gives you a score that moves, scored against six objective checks, with errors caught on Wednesday by the person signing on Friday.

2 · Your sample finds nine of twelve reports are missing the ask. What do you do first?

Review the artifact, never the person, and fix the system before the people. Nine out of twelve is a template problem. Publishing individual scores would also empty every gaps block in the function within a month.

3 · Your director asks for one number at week 4. Which of these is worth giving them?

Adoption rate, prompts and seats measure logging in, not delivery, and any PM could move them in an afternoon. Score plus hours, quoted together and as a shape rather than false precision, is the pair nobody can quote half of.

Leader session 6 cheat sheet · pin this

The pilot shapeone programme, one artifact, six weeks, one PM who volunteered. Nothing else changes.
Why the status reporthighest frequency, clearest rubric, lowest blast radius. Boring beats exciting.
Baseline in week 1three reports scored, one hours estimate, three slips counted. Twenty minutes, and it makes the case.
Sample, do not police1 in 5 at random. Artifact reviewed, never the person. Aggregate published, individuals never.
Fix the system firsta check that fails across four PMs is a template bug, not a competence problem.
Honest metricshours moved · report score · recurring gaps · decisions per steerco · slip-to-escalation days.
Vanity metricsprompts, tokens, adoption %, seats, unbaselined time saved. All gameable in an afternoon.
Say it out loudopt-in, never a leaderboard, stopping is a result, and nobody's accountability moved.