learn-ai-project-management-with-phoebe / Leader session 4 of 6
Learn AI + Project Management with Phoebe · Leader track · Session 4 of 6

Risk, dependencies and escalation: the signals worth waking up for

Point AI at a delivery week and it will surface twenty things. Two of them matter. Tonight you learn which two, why a hard gate is the one thing a cheerful summariser will always smooth over, and how to build an escalation that survives a steerco: the decision, the options with their real costs, your recommendation, and the date after which the option expires.

🟡 Leader track Heads of delivery Sponsors Escalation drill 45 min
0-3 · Where we are 3-20 · Signals, hard gates 20-42 · Triage and rewrite drills 42-45 · Q&A
Part 0

Where we are

a1 drew the line between drafting and judgment. a2 built the governed setup. a3 gave you a seven-point rubric for the plans you sign, and ended on the finding that always comes next: a dependency nobody confirmed is not a plan defect, it is an open risk. Tonight we pick that up. Northwind is in week nine, its M3 gate has moved from 10 August to a 24 August forecast, and the question in front of the sponsor is not "what happened" - it is "what do you want me to decide, and by when".

Live - presented in session Self-study - read after class ★ Try it now prompt Official sources covered
★ What you walk out with today A working filter for which AI-surfaced early warnings to trust and which to kill on sight, the ability to tell a soft dependency from a hard gate in someone else's cheerful summary, the four-part anatomy of an escalation a sponsor can act on in ninety seconds, and a rewritten Northwind escalation you could send tomorrow.
Part 1 · covers PMBOK 7 uncertainty domain

Early warnings: what to trust, what to kill 7 min live

Early-warning detection is one of the three places AI lands first in delivery, and it is the one most likely to be oversold. The honest version is narrow and genuinely useful: a model with your risk register, your plan and your last four reports is excellent at noticing what changed. It is worthless at predicting what people will do. Those two capabilities look identical on the page, which is the whole problem.

WHAT A MACHINE CAN VERIFY WHAT ONLY THE ROOM KNOWS Signal - keep it a risk whose trigger has actually fired the same failure three nights running an action that has rolled for five weeks this week diffed against last week Noise - kill it a probability the model made up a two-point wobble called a trend sentiment read off meeting notes any guess about who might resign AI is good at diffing, counting and remembering. It is bad at guessing what people will do. Keep what a machine can check. Kill what only a human in the room could know. The best AI early warning is boring: this changed, that did not, this has been open 11 days.
🔍 Click to zoom - keep the verifiable, kill the invented; the useful signals are the dull ones
LiveThe three things it is genuinely good at3 min
  • Diffing this week against last week. Humans are terrible at this. You read the report, it feels familiar, and you miss that the DQ pass rate has been quietly flat at 96.4% for three weeks while the commentary has said "improving" every time. A machine holding four reports does not get bored.
  • Spotting a risk whose trigger has fired. Your register says R-11 triggers if DQ sits below the 99% gate for two consecutive weeks. That is a rule. A model can check it against the numbers every Monday without anyone remembering to.
  • Noticing the action that has been rolling. "Chase Veridian" has moved forward five weeks in a row, in five different reports, and nobody noticed because each individual roll looked reasonable. Age is a fact. It is also the single most reliable predictor of an action that is never going to happen.
Real world

Five weeks of "next week". A delivery lead asked her workspace for every action that had moved more than twice. Nine came back. Eight were noise. The ninth was ticket VN-2291 - chase Veridian on the POS extract SLA - which had been rolled five times and was now eleven days open with the vendor. The delay was not the problem. The problem was that five weeks of rolling had turned an unconfirmed dependency into a live risk without a single person deciding to let it happen.

LiveThe three things it is genuinely bad at3 min
  • Probability. Ask for the likelihood of a risk and you will get a number. It will be confident, it will be plausible, and it is generated, not calculated. There is no data behind "68%". A generated probability is worse than no probability, because it survives the trip to the steerco slide where nobody can interrogate it.
  • Knowing which vendor will deliver. That knowledge lives in Priya N.'s head, built from four years of watching Veridian miss soft deadlines and hit contractual ones. No document in your workspace contains it. The model can tell you the SLA says 5 days and the ticket is 11 days old - facts. It cannot tell you whether this vendor is about to come good.
  • People. It cannot tell you Marcus is exhausted, that Sofia is interviewing elsewhere, or that Dan has quietly stopped believing in the M5 date. Sentiment analysis on meeting notes is the most seductive false signal in this whole space: it produces a chart, and the chart is measuring how someone writes, not how they feel.
The one-question filter For any AI-surfaced warning ask: could I verify this from a document or a system in under a minute? Trigger fired, ticket age, count of failed runs, gate distance - yes. Likelihood, mood, intent, whether they will make it - no. Keep the first kind, kill the second, and never let the second kind reach a slide.
Self-studyWhy the noise is so convincing2 min read

Because it is written in the same register as the signal. "R-07 likelihood has increased to high given deteriorating vendor engagement" and "VN-2291 has been open 11 days against a 5-day SLA" look equally authoritative. One is an assertion with nothing underneath it, the other is a fact you can click, and nothing in the prose tells them apart. PMBOK's uncertainty domain separates what you can characterise from what you can only prepare for: AI is strong on characterisation - what is true now, what changed, what threshold has been crossed - and adds nothing to preparation, which is judgment, options and appetite. PMI's 2026 AI standard carries risk as one of its eight equally weighted principles, paired with governance and data quality, for exactly this reason: a confident output built on thin data is itself the risk, and the guardrail is human oversight with real intervention triggers rather than a review that rubber-stamps whatever arrived.

Part 2 · the failure mode that stops programmes

The hard gates AI keeps smoothing over 6 min live

A soft dependency slips and the programme absorbs it. A hard gate stops the programme dead. The difference is enormous, it is obvious once someone says it out loud, and it is invisible in a fluent summary - because a model summarising cheerfully will describe both of them in the same reasonable tone.

LiveNorthwind R-03: the plan that reads beautifully and cannot happen3 min

Northwind's risk register has R-03: data-engineer hiring gap. It is not a delay risk. It is a hard gate before M4, because you cannot build a modelled warehouse with nobody to build it - no compression, no partial credit, no working around it with overtime. Zero engineers produces zero model, at any date you choose to write down. Now read how the draft plan handles it: "M4 build begins immediately on M3 sign-off with no gap." That sentence is not wrong about the sequence. It is silent about the requirement. The plan flows, the dates line up, and the assumption underneath it - two people who have not been hired - never appears on the page at all.

Soft dependencyHard gate
ExampleThe glossary sign-off is a week lateR-03: no data engineer before M4
What slipping costsAbsorbed, or float is consumedThe next milestone cannot start at all
Recovery leversResequence, parallelise, trim scopeOnly one: remove the constraint
How a draft describes it"minor slip, being managed""minor slip, being managed"
Who must knowThe delivery leadThe sponsor, immediately

Row four is the point of this session. Both read identically in the summary. You cannot detect the difference by reading better, so you detect it by asking a structural question instead: for each item on this list, is there any version of the next milestone that starts without it? If the answer is no, it is a gate, and it goes to the sponsor this week rather than next.

Real world

Six weeks of green. A programme reported amber-then-green for six weeks on a workstream whose entire delivery depended on two contractors nobody had approved. Every weekly summary was accurate: design was progressing, the environment was ready, the backlog was groomed. All true, all irrelevant. When the build date arrived there was nobody to build, and the sponsor's first question was the right one - "when did you know?" The honest answer was week two. The summary had simply never been asked whether anything on it was a gate.

LiveMaking a model surface gates instead of smoothing them2 min

The fix is a standing instruction in your workspace, not a better weekly prompt. Cheerful smoothing is a default, and defaults are changed once.

★ Try it now - the anti-smoothing rule for your house rulesHARD GATES (apply to every plan, report or risk summary you draft) Before summarising status, separate items into two lists: BLOCKING the next milestone cannot start without it at all - missing people, environment, approval, data or signature. NON-BLOCKING everything else, including anything that can be resequenced, compressed, parallelised or absorbed. For every BLOCKING item give one line: what is missing, who owns removing it, the last date it can go without moving the next milestone, and whether that date has passed. Never soften a blocking item. Not "minor", not "being managed", not "on track to resolve". If it is unresolved, say unresolved and give the date. If the plan assumes a person, a budget or an approval that is not in the pack, name it as an assumption before anything else.
The sponsor's version of the same question Two things make the rule work: it defines blocking structurally, so the model classifies rather than judges, and it bans the softening vocabulary outright, because "being managed" is exactly the phrase that lets a gate travel through four reports unnoticed. If you sit on the sponsor's side of the table you need only one line: "which of these can we not start without?" It cuts through any summary, AI-drafted or not.
Part 3 · the artifact that has to survive a room

The escalation that survives a steerco 6 min live

Most escalations fail for the same reason: they report a situation instead of requesting a decision. The sponsor reads it, feels concerned, asks two questions nobody can answer, and the item rolls to next month - by which point the option that mattered has expired. Four parts fix this permanently.

FOUR PARTS, IN THIS ORDER, ON ONE PAGE The decision one sentence, so a yes or no ends it The options two or three, each with its real cost in money, time or scope Your pick named, with the reason in one line - never a menu alone The expiry the date after which this option is gone, and what goes then An escalation with no decision is a complaint. With no expiry it is a complaint with data. Steercos do not fail because the news is bad. They fail because nothing can be decided. Elena V. can say yes in ninety seconds when all four parts are there. That is the bar.
🔍 Click to zoom - decision, costed options, your recommendation, expiry date; all four or it rolls
LiveThe Northwind M3 case, in four parts3 min
  • The position. Week nine. M3 was gated 10 August, forecast is now 24 August. Six of nine sources landing at 812k rows a night, DQ pass rate 96.4% against a 99% gate, NWD-412 failed the CDC load three consecutive nights. Behind it, R-03 means M4 has nobody to build it.
  • Decision. Approve the M3 gate move to 24 August, or fund two contract data engineers for four weeks to hold the original ladder.
  • Options with real costs. Option A: move the gate to 24 August. Cost - M4 and everything behind it moves two weeks, and D-15 keeps the legacy warehouse running in parallel for longer. Option B: fund two contract data engineers for four weeks at approximately £38k, holding M4's start. Cost - the money, plus onboarding time Marcus L. does not currently have.
  • Recommendation. Yours, named, with one reason. Not a menu handed to the sponsor to sort out.
  • Expiry. The decision is needed by 8 August. After that, contractor lead time means option B cannot land before M4 anyway, and M4 slips whichever way the sponsor votes. This is the part people leave out and the part that changes behaviour: "we need a decision" gets read, "we need a decision by 8 August or the second option stops existing" gets decided.
  • What AI may and may not touch here. It can assemble the evidence - ticket ids, dates, gate distance, decision log lines - force the four parts, and list the questions Elena V. will ask. It must not write the recommendation, which carries your name, or the costs: the £38k comes from procurement, the four weeks from Marcus, the lead time from the contract. Never from a draft.
Real world

The escalation that finally moved. A delivery lead had raised the same resourcing concern in four consecutive steercos. Each time it was noted, each time it rolled. The fifth time she wrote two options with numbers against them and one line: "after the 8th, option B cannot land in time and this becomes a slip, not a choice." The sponsor approved in under two minutes. Nothing about the underlying problem had changed. What changed was that for the first time there was something to say yes to, and a reason not to say it next month.

Self-studyWhy the escalation is where rail three bites hardest2 min read

A named human signs it. PMI's standard puts human oversight with real intervention triggers at the centre of AI-assisted delivery, and the escalation is the single artifact where that matters most, because it is the one that ends in somebody spending money or accepting a slip. A status report that is slightly wrong gets corrected next week. An escalation that is slightly wrong gets acted on. So the collation can be drafted, the shape can be enforced by a rule, and the two things that decide the outcome - the recommendation and the numbers - come from people with names, every time.

Demo 1 of 2

Triage six early warnings: keep two, kill four ★ 12 min · everyone triages

Northwind's Monday scan surfaced these six. Exactly two belong in front of a human this week. Four should be killed, and you should be able to say why in one sentence each. Four minutes on your own, then compare.

W1. "The plan has M4 build starting immediately on M3 sign-off. No named data engineer is assigned to M4 in the owner list, and no requisition appears in the pack. Risk R-03 is open."

W2. "R-07 vendor SLA: likelihood has increased to 72% based on deteriorating engagement signals across recent correspondence."

W3. "Ticket NWD-412 records a failed CDC load on three consecutive nights. Nightly volume is 812k rows and 6 of 9 sources are landing."

W4. "Open actions rose from 9 last week to 12 this week, indicating workload pressure on the delivery team."

W5. "DQ pass rate moved from 96.6% to 96.4% week on week, continuing a downward trend."

W6. "Sentiment across the last three stand-up notes is more negative than the preceding three weeks, suggesting declining team morale."

LiveThe triage answers - open after you have made your calls4 min
CallWhy
W1KEEPA hard gate stated as a fact about the pack, verifiable in a minute. R-03 has no owner assigned to M4 and the plan assumes people who do not exist. This is the escalation.
W2Kill72% is generated, not calculated. There is no data behind it. The verifiable version is the one to keep instead: VN-2291 has been open 11 days against a 5-day SLA. Ask Priya N. for the judgment.
W3KEEPThree consecutive failures is a countable threshold, and it is the mechanism behind the M3 slip. It also tells you the 24 August forecast has a cause, not a hope, behind it.
W4KillA count is not a signal. Twelve fresh actions is a healthy week; three actions rolled five times is a problem. Ask for age and owner, not volume. The useful version of W4 is "which actions have moved more than twice".
W5KillTwo points is not a trend, and 0.2 points is inside the noise. Distance to the 99% gate has not meaningfully changed. Keep watching it, do not surface it - and note that R-11 triggers on a rule, which has not fired.
W6KillSentiment on meeting notes measures how somebody writes, not how the team feels. If you are worried about morale, the instrument is a conversation with Marcus L., not a chart.

The pattern in the four kills: two invented a number (W2, W5), one counted the wrong thing (W4), and one measured a proxy for something it cannot see (W6). Every one would have looked perfectly respectable on a slide.

Demo 2 of 2

Rewrite a weak escalation into a defensible one ★ 10 min · build your own

Here is the escalation Northwind's draft pack actually produced. Read it, then rebuild it using the four parts before you look at the version below. Ten minutes, on paper, alone or in pairs. The strong version is not longer because it says more - it is longer because it stopped hiding.

The weak version - what actually got draftedESCALATION - Northwind data platform, week 9 We are concerned about the vendor and may need more resource. The ingest workstream is experiencing some challenges and M3 is now expected to be later than originally planned. The team is working hard to recover the position and we are managing the risks closely. We would welcome the steerco's support.

Find the decision. There is not one. Write the single sentence the sponsor is being asked to say yes or no to.

Find the options. There are none. Write two, and put a real cost against each - money, time or scope. Take the numbers from the canon, not from your imagination.

Find the recommendation. Absent, and "we would welcome support" is not one. Pick an option and give a one-line reason.

Find the expiry. Absent, which is why this item will roll. Write the date after which one of the options stops existing, and what happens then. Then delete every softening word - "some challenges", "working hard", "managing closely", "concerned" - because each one is doing the job a number should be doing.

★ The defensible version - all four parts, one pageESCALATION - Northwind data platform, week 9 For decision at steerco, 6 Aug. Owner: delivery lead. Sponsor: Elena V. DECISION NEEDED Approve the M3 gate move from 10 Aug to 24 Aug, OR fund two contract data engineers for four weeks to hold the original M4 start. POSITION 6 of 9 sources landing, 812k rows/night. DQ 96.4% against a 99% gate. NWD-412: CDC load failed 3 consecutive nights. M3 forecast 24 Aug, sourced from Marcus L. and nightly throughput. VN-2291 open 11 days with Veridian (R-07). Hard gate behind it: R-03, nobody assigned to M4. OPTIONS A. Move the M3 gate to 24 Aug. Cost: M4 and all after it move ~2 weeks; legacy warehouse runs parallel longer (D-15). No new spend. B. Fund 2 contract data engineers, 4 weeks, approx GBP 38k. Cost: the spend plus onboarding time Marcus L. does not have, and it holds the M4 start only if approved in time. RECOMMENDATION Option A. The 24 Aug forecast is evidence-based and contractors cannot be productive before M4 starts, so B buys risk rather than time. DECISION BY 8 AUG After 8 Aug, contractor lead time means option B cannot land before M4: this stops being a choice and becomes a slip. Escalate to Elena V. direct if the steerco does not resolve it on 6 Aug.
Real world

"Concerned" is not a status. A sponsor kept a list of words she sent straight back: concerned, challenges, closely, shortly, broadly. Her rule was that each one is a number wearing a coat. It sounded pedantic for about a month, and then the escalations arriving at her steerco started containing dates and costs, and the meetings got forty minutes shorter. She had not asked for better writing. She had made vagueness cost more than precision.

Homework

Try it yourself - this week ◐ 30-45 min total

Source material

Official sources covered

Taught from the delivery canon and PMI's public AI guidance. Certification (PMP, PMI-ACP) and the full normative text of the AI standard stay with PMI. This session covers:

PMBOK Guide 7th ed. - uncertainty performance domainParts 1-3 · characterising risk, options, response and appetite
PMI - Standard for AI in Portfolio, Program and Project Management (2026)Parts 1-3 · the risk principle, oversight with real intervention triggers
This course, PM track b4 - dependencies, gates and the critical pathPart 2 · finding the hard gate in the dependency map
This course, PM track b5 - a risk register that maintains itselfPart 1 · triggers, ageing, and the weekly scan in working depth
Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · A weekly scan reports "R-07 likelihood has increased to 72%". What do you do with it?

A generated probability has no data behind it, and it survives to the slide precisely because it looks rigorous. Keep what a machine can verify - ticket age, gate distance, failed run counts - and get the likelihood judgment from Priya N., who knows the vendor.

2 · Northwind's summary says R-03 is "a minor resourcing slip, being managed". Why is that dangerous?

A soft dependency slips and is absorbed. A hard gate stops the programme. A cheerful summariser describes both in the same reasonable tone, so use the structural test instead: is there any version of the next milestone that starts without this?

3 · Which part is missing from "we are concerned about the vendor and may need more resource"?

It reports a feeling, not a request. A defensible escalation names the decision, gives two or three options with real costs, carries your recommendation, and states the date after which an option expires. Without the expiry it rolls to next month, by which time it is a slip rather than a choice.

Leader session 4 cheat sheet · pin this

Good atdiffing week on week, triggers that have fired, actions that have rolled. Memory and comparison.
Bad atprobability, which vendor will deliver, and anything about how a person feels or what they will do.
The one-question filterCould I verify this from a document or system in under a minute? If not, kill it.
Hard gate testIs there any version of the next milestone that starts without this? No means gate, and the sponsor hears it this week.
R-03 in one lineYou cannot model a warehouse with nobody to build it. No compression, no partial credit.
Banned vocabularyconcerned, challenges, closely, being managed, working hard. Each one is a number wearing a coat.
Escalation, four partsthe decision · two or three costed options · your named recommendation · the expiry date.
The expiry line"After 8 Aug this stops being a choice and becomes a slip." That is what gets it decided.