Where we are
a1 drew the line between drafting and judgment. a2 built the governed setup. a3 gave you a seven-point rubric for the plans you sign, and ended on the finding that always comes next: a dependency nobody confirmed is not a plan defect, it is an open risk. Tonight we pick that up. Northwind is in week nine, its M3 gate has moved from 10 August to a 24 August forecast, and the question in front of the sponsor is not "what happened" - it is "what do you want me to decide, and by when".
Early warnings: what to trust, what to kill 7 min live
Early-warning detection is one of the three places AI lands first in delivery, and it is the one most likely to be oversold. The honest version is narrow and genuinely useful: a model with your risk register, your plan and your last four reports is excellent at noticing what changed. It is worthless at predicting what people will do. Those two capabilities look identical on the page, which is the whole problem.
LiveThe three things it is genuinely good at3 min▶
- Diffing this week against last week. Humans are terrible at this. You read the report, it feels familiar, and you miss that the DQ pass rate has been quietly flat at 96.4% for three weeks while the commentary has said "improving" every time. A machine holding four reports does not get bored.
- Spotting a risk whose trigger has fired. Your register says R-11 triggers if DQ sits below the 99% gate for two consecutive weeks. That is a rule. A model can check it against the numbers every Monday without anyone remembering to.
- Noticing the action that has been rolling. "Chase Veridian" has moved forward five weeks in a row, in five different reports, and nobody noticed because each individual roll looked reasonable. Age is a fact. It is also the single most reliable predictor of an action that is never going to happen.
Five weeks of "next week". A delivery lead asked her workspace for every action that had moved more than twice. Nine came back. Eight were noise. The ninth was ticket VN-2291 - chase Veridian on the POS extract SLA - which had been rolled five times and was now eleven days open with the vendor. The delay was not the problem. The problem was that five weeks of rolling had turned an unconfirmed dependency into a live risk without a single person deciding to let it happen.
LiveThe three things it is genuinely bad at3 min▶
- Probability. Ask for the likelihood of a risk and you will get a number. It will be confident, it will be plausible, and it is generated, not calculated. There is no data behind "68%". A generated probability is worse than no probability, because it survives the trip to the steerco slide where nobody can interrogate it.
- Knowing which vendor will deliver. That knowledge lives in Priya N.'s head, built from four years of watching Veridian miss soft deadlines and hit contractual ones. No document in your workspace contains it. The model can tell you the SLA says 5 days and the ticket is 11 days old - facts. It cannot tell you whether this vendor is about to come good.
- People. It cannot tell you Marcus is exhausted, that Sofia is interviewing elsewhere, or that Dan has quietly stopped believing in the M5 date. Sentiment analysis on meeting notes is the most seductive false signal in this whole space: it produces a chart, and the chart is measuring how someone writes, not how they feel.
Self-studyWhy the noise is so convincing2 min read▶
Because it is written in the same register as the signal. "R-07 likelihood has increased to high given deteriorating vendor engagement" and "VN-2291 has been open 11 days against a 5-day SLA" look equally authoritative. One is an assertion with nothing underneath it, the other is a fact you can click, and nothing in the prose tells them apart. PMBOK's uncertainty domain separates what you can characterise from what you can only prepare for: AI is strong on characterisation - what is true now, what changed, what threshold has been crossed - and adds nothing to preparation, which is judgment, options and appetite. PMI's 2026 AI standard carries risk as one of its eight equally weighted principles, paired with governance and data quality, for exactly this reason: a confident output built on thin data is itself the risk, and the guardrail is human oversight with real intervention triggers rather than a review that rubber-stamps whatever arrived.
The hard gates AI keeps smoothing over 6 min live
A soft dependency slips and the programme absorbs it. A hard gate stops the programme dead. The difference is enormous, it is obvious once someone says it out loud, and it is invisible in a fluent summary - because a model summarising cheerfully will describe both of them in the same reasonable tone.
LiveNorthwind R-03: the plan that reads beautifully and cannot happen3 min▶
Northwind's risk register has R-03: data-engineer hiring gap. It is not a delay risk. It is a hard gate before M4, because you cannot build a modelled warehouse with nobody to build it - no compression, no partial credit, no working around it with overtime. Zero engineers produces zero model, at any date you choose to write down. Now read how the draft plan handles it: "M4 build begins immediately on M3 sign-off with no gap." That sentence is not wrong about the sequence. It is silent about the requirement. The plan flows, the dates line up, and the assumption underneath it - two people who have not been hired - never appears on the page at all.
| Soft dependency | Hard gate | |
|---|---|---|
| Example | The glossary sign-off is a week late | R-03: no data engineer before M4 |
| What slipping costs | Absorbed, or float is consumed | The next milestone cannot start at all |
| Recovery levers | Resequence, parallelise, trim scope | Only one: remove the constraint |
| How a draft describes it | "minor slip, being managed" | "minor slip, being managed" |
| Who must know | The delivery lead | The sponsor, immediately |
Row four is the point of this session. Both read identically in the summary. You cannot detect the difference by reading better, so you detect it by asking a structural question instead: for each item on this list, is there any version of the next milestone that starts without it? If the answer is no, it is a gate, and it goes to the sponsor this week rather than next.
Six weeks of green. A programme reported amber-then-green for six weeks on a workstream whose entire delivery depended on two contractors nobody had approved. Every weekly summary was accurate: design was progressing, the environment was ready, the backlog was groomed. All true, all irrelevant. When the build date arrived there was nobody to build, and the sponsor's first question was the right one - "when did you know?" The honest answer was week two. The summary had simply never been asked whether anything on it was a gate.
LiveMaking a model surface gates instead of smoothing them2 min▶
The fix is a standing instruction in your workspace, not a better weekly prompt. Cheerful smoothing is a default, and defaults are changed once.
The escalation that survives a steerco 6 min live
Most escalations fail for the same reason: they report a situation instead of requesting a decision. The sponsor reads it, feels concerned, asks two questions nobody can answer, and the item rolls to next month - by which point the option that mattered has expired. Four parts fix this permanently.
LiveThe Northwind M3 case, in four parts3 min▶
- The position. Week nine. M3 was gated 10 August, forecast is now 24 August. Six of nine sources landing at 812k rows a night, DQ pass rate 96.4% against a 99% gate, NWD-412 failed the CDC load three consecutive nights. Behind it, R-03 means M4 has nobody to build it.
- Decision. Approve the M3 gate move to 24 August, or fund two contract data engineers for four weeks to hold the original ladder.
- Options with real costs. Option A: move the gate to 24 August. Cost - M4 and everything behind it moves two weeks, and D-15 keeps the legacy warehouse running in parallel for longer. Option B: fund two contract data engineers for four weeks at approximately £38k, holding M4's start. Cost - the money, plus onboarding time Marcus L. does not currently have.
- Recommendation. Yours, named, with one reason. Not a menu handed to the sponsor to sort out.
- Expiry. The decision is needed by 8 August. After that, contractor lead time means option B cannot land before M4 anyway, and M4 slips whichever way the sponsor votes. This is the part people leave out and the part that changes behaviour: "we need a decision" gets read, "we need a decision by 8 August or the second option stops existing" gets decided.
- What AI may and may not touch here. It can assemble the evidence - ticket ids, dates, gate distance, decision log lines - force the four parts, and list the questions Elena V. will ask. It must not write the recommendation, which carries your name, or the costs: the £38k comes from procurement, the four weeks from Marcus, the lead time from the contract. Never from a draft.
The escalation that finally moved. A delivery lead had raised the same resourcing concern in four consecutive steercos. Each time it was noted, each time it rolled. The fifth time she wrote two options with numbers against them and one line: "after the 8th, option B cannot land in time and this becomes a slip, not a choice." The sponsor approved in under two minutes. Nothing about the underlying problem had changed. What changed was that for the first time there was something to say yes to, and a reason not to say it next month.
Self-studyWhy the escalation is where rail three bites hardest2 min read▶
A named human signs it. PMI's standard puts human oversight with real intervention triggers at the centre of AI-assisted delivery, and the escalation is the single artifact where that matters most, because it is the one that ends in somebody spending money or accepting a slip. A status report that is slightly wrong gets corrected next week. An escalation that is slightly wrong gets acted on. So the collation can be drafted, the shape can be enforced by a rule, and the two things that decide the outcome - the recommendation and the numbers - come from people with names, every time.
Triage six early warnings: keep two, kill four ★ 12 min · everyone triages
Northwind's Monday scan surfaced these six. Exactly two belong in front of a human this week. Four should be killed, and you should be able to say why in one sentence each. Four minutes on your own, then compare.
W1. "The plan has M4 build starting immediately on M3 sign-off. No named data engineer is assigned to M4 in the owner list, and no requisition appears in the pack. Risk R-03 is open."
W2. "R-07 vendor SLA: likelihood has increased to 72% based on deteriorating engagement signals across recent correspondence."
W3. "Ticket NWD-412 records a failed CDC load on three consecutive nights. Nightly volume is 812k rows and 6 of 9 sources are landing."
W4. "Open actions rose from 9 last week to 12 this week, indicating workload pressure on the delivery team."
W5. "DQ pass rate moved from 96.6% to 96.4% week on week, continuing a downward trend."
W6. "Sentiment across the last three stand-up notes is more negative than the preceding three weeks, suggesting declining team morale."
LiveThe triage answers - open after you have made your calls4 min▶
| Call | Why | |
|---|---|---|
| W1 | KEEP | A hard gate stated as a fact about the pack, verifiable in a minute. R-03 has no owner assigned to M4 and the plan assumes people who do not exist. This is the escalation. |
| W2 | Kill | 72% is generated, not calculated. There is no data behind it. The verifiable version is the one to keep instead: VN-2291 has been open 11 days against a 5-day SLA. Ask Priya N. for the judgment. |
| W3 | KEEP | Three consecutive failures is a countable threshold, and it is the mechanism behind the M3 slip. It also tells you the 24 August forecast has a cause, not a hope, behind it. |
| W4 | Kill | A count is not a signal. Twelve fresh actions is a healthy week; three actions rolled five times is a problem. Ask for age and owner, not volume. The useful version of W4 is "which actions have moved more than twice". |
| W5 | Kill | Two points is not a trend, and 0.2 points is inside the noise. Distance to the 99% gate has not meaningfully changed. Keep watching it, do not surface it - and note that R-11 triggers on a rule, which has not fired. |
| W6 | Kill | Sentiment on meeting notes measures how somebody writes, not how the team feels. If you are worried about morale, the instrument is a conversation with Marcus L., not a chart. |
The pattern in the four kills: two invented a number (W2, W5), one counted the wrong thing (W4), and one measured a proxy for something it cannot see (W6). Every one would have looked perfectly respectable on a slide.
Rewrite a weak escalation into a defensible one ★ 10 min · build your own
Here is the escalation Northwind's draft pack actually produced. Read it, then rebuild it using the four parts before you look at the version below. Ten minutes, on paper, alone or in pairs. The strong version is not longer because it says more - it is longer because it stopped hiding.
Find the decision. There is not one. Write the single sentence the sponsor is being asked to say yes or no to.
Find the options. There are none. Write two, and put a real cost against each - money, time or scope. Take the numbers from the canon, not from your imagination.
Find the recommendation. Absent, and "we would welcome support" is not one. Pick an option and give a one-line reason.
Find the expiry. Absent, which is why this item will roll. Write the date after which one of the options stops existing, and what happens then. Then delete every softening word - "some challenges", "working hard", "managing closely", "concerned" - because each one is doing the job a number should be doing.
"Concerned" is not a status. A sponsor kept a list of words she sent straight back: concerned, challenges, closely, shortly, broadly. Her rule was that each one is a number wearing a coat. It sounded pedantic for about a month, and then the escalations arriving at her steerco started containing dates and costs, and the meetings got forty minutes shorter. She had not asked for better writing. She had made vagueness cost more than precision.
Try it yourself - this week ◐ 30-45 min total
- Take your live risk register and mark every item blocking or non-blocking using the one test: is there any version of the next milestone that starts without it? Count how many blocking items your last status report described as "being managed".
- Ask your team for every action that has moved more than twice. Not the action count - the age. Expect one uncomfortable answer, and find one number on a current slide that came from a model rather than a person or a system. Delete it or source it. There is usually one.
- Add the anti-smoothing rule to your PMO's house rules and write the banned words down, then rewrite your most recent real escalation into the four parts. If you cannot find an expiry date for it, that tells you why it rolled.
Official sources covered
Taught from the delivery canon and PMI's public AI guidance. Certification (PMP, PMI-ACP) and the full normative text of the AI standard stay with PMI. This session covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · A weekly scan reports "R-07 likelihood has increased to 72%". What do you do with it?
A generated probability has no data behind it, and it survives to the slide precisely because it looks rigorous. Keep what a machine can verify - ticket age, gate distance, failed run counts - and get the likelihood judgment from Priya N., who knows the vendor.
2 · Northwind's summary says R-03 is "a minor resourcing slip, being managed". Why is that dangerous?
A soft dependency slips and is absorbed. A hard gate stops the programme. A cheerful summariser describes both in the same reasonable tone, so use the structural test instead: is there any version of the next milestone that starts without this?
3 · Which part is missing from "we are concerned about the vendor and may need more resource"?
It reports a feeling, not a request. A defensible escalation names the decision, gives two or three options with real costs, carries your recommendation, and states the date after which an option expires. Without the expiry it rolls to next month, by which time it is a slip rather than a choice.